FWPRec Experiments
All experiments
Table 2 campaign · seeds 42/43/44 · updated 2026-09-23 11:55 UTC

The paper's table, run to spec

This run treats the paper's own text as law and fills its main comparison table from nothing: no number, checkpoint, or shortcut from any earlier campaign. Every model reads the same frozen 64-dimensional text vectors for items, has its learned per-item bias removed, sees identical splits and candidates, and is picked by validation score before touching the test set once. Where the paper names training settings we use them verbatim; where it is silent, each model earned its learning rate through the same three-way validation screen. The one licensed change: MIND-large stands in for MIND-small, run last.

So far

Four of five columns are final, zero failed runs. The notepad (FWPRec) leads Amazon and EB-NeRD outright, sits mid-table on MovieLens-1M where plain item co-occurrence still rules, and is in a photo finish with SASRec on KuaiRec while holding the best AUC there. MIND-large is training now.

The deal every model gets

Same frozen Qwen3-Embedding-0.6B item vectors (64-d, length-normalized) · no learned item bias · same chronological splits, histories, candidates and labels (construction seed 42) · positive-only histories for sequence models, with the signed stream reaching only the notepad's memory · best-validation checkpoint, then one pass over test · means over training seeds 42/43/44. Popularity and ItemKNN use no item vectors at all — they are the representation-independent floor, run once. BPRMF with a frozen item table trains only its user vectors, a real capacity cut, noted here so the row is read fairly.

MIND-large

News clicks at full scale (~376k test impressions). The paper's MIND-small settings applied unchanged, per the owner's substitution instruction.

⏳ In progress — 3 of 10 model rows final. Numbers shown are complete three-seed means; missing rows are still training.

ModelAUCMRRnDCG@5nDCG@10seeds
Popularity0.49410.25270.22530.29101/1
ItemKNN0.50340.26430.23660.30281/1
BPRMF0.5771 ±0.00050.3223 ±0.00030.2974 ±0.00030.3635 ±0.00043/3
GRU4Rec0.61760.34990.32780.39601/3 ⏳
SASRec0.57100.31380.28620.35841/3 ⏳
DuoRec0.6152 ±0.00010.3554 ±0.00170.3327 ±0.00130.3983 ±0.00112/3 ⏳
LRURec····running
Mamba4Rec0.6197 ±0.00280.3536 ±0.00100.3320 ±0.00150.3999 ±0.00172/3 ⏳
HSTU0.62630.35920.33770.40501/3 ⏳
FWPRec (Ours)····running

EB-NeRD

Danish news clicks. The officially labeled validation partition serves as the held-out test, so this column is test numbers like every other.

ModelAUCMRRnDCG@5nDCG@10seeds
Popularity0.44360.26350.28930.38541/1
ItemKNN0.48940.30140.33300.41971/1
BPRMF0.5623 ±0.00070.3612 ±0.00050.3997 ±0.00050.4749 ±0.00053/3
GRU4Rec0.5633 ±0.00080.3586 ±0.00050.3976 ±0.00070.4729 ±0.00073/3
SASRec0.5328 ±0.00740.3388 ±0.00640.3747 ±0.00690.4544 ±0.00593/3
DuoRec0.5424 ±0.00660.3422 ±0.00570.3792 ±0.00650.4577 ±0.00523/3
LRURec0.5635 ±0.00100.3586 ±0.00060.3976 ±0.00080.4731 ±0.00083/3
Mamba4Rec0.5640 ±0.00160.3594 ±0.00120.3984 ±0.00120.4736 ±0.00123/3
HSTU0.5620 ±0.00050.3572 ±0.00090.3963 ±0.00080.4717 ±0.00073/3
FWPRec (Ours)0.5699 ±0.00440.3642 ±0.00310.4043 ±0.00370.4777 ±0.00313/3

Amazon Video Games

Product reviews with star ratings. Test set: 19,773 impressions, the paper's exact number.

ModelAUCMRRnDCG@5nDCG@10seeds
Popularity0.31260.07070.05980.07451/1
ItemKNN0.53200.09650.07890.09531/1
BPRMF0.5572 ±0.00130.0761 ±0.00070.0553 ±0.00080.0769 ±0.00113/3
GRU4Rec0.7957 ±0.00270.1910 ±0.00310.1790 ±0.00340.2292 ±0.00363/3
SASRec0.7718 ±0.00090.1823 ±0.00110.1713 ±0.00170.2153 ±0.00183/3
DuoRec0.7793 ±0.00280.1621 ±0.00210.1456 ±0.00230.1935 ±0.00363/3
LRURec0.8114 ±0.00170.2145 ±0.00210.2075 ±0.00220.2566 ±0.00263/3
Mamba4Rec0.8055 ±0.00170.2046 ±0.00020.1951 ±0.00050.2442 ±0.00073/3
HSTU0.8146 ±0.00380.2247 ±0.00260.2184 ±0.00320.2674 ±0.00313/3
FWPRec (Ours)0.8233 ±0.00200.2390 ±0.00200.2342 ±0.00280.2825 ±0.00233/3

MovieLens-1M

Movie ratings, split 80/10/10 inside each user's own timeline (the paper's wording, implemented fresh for this run).

ModelAUCMRRnDCG@5nDCG@10seeds
Popularity0.82860.20630.19480.24781/1
ItemKNN0.85570.29460.30350.36401/1
BPRMF0.6631 ±0.00000.1043 ±0.00010.0826 ±0.00000.1150 ±0.00013/3
GRU4Rec0.7997 ±0.00040.2372 ±0.00040.2356 ±0.00040.2833 ±0.00063/3
SASRec0.7909 ±0.00070.2408 ±0.00050.2392 ±0.00050.2848 ±0.00083/3
DuoRec0.7988 ±0.00030.2560 ±0.00210.2565 ±0.00240.3015 ±0.00203/3
LRURec0.8112 ±0.00030.2816 ±0.00090.2842 ±0.00090.3278 ±0.00103/3
Mamba4Rec0.8102 ±0.00110.2803 ±0.00100.2828 ±0.00090.3264 ±0.00133/3
HSTU0.8159 ±0.00080.2914 ±0.00120.2949 ±0.00170.3378 ±0.00123/3
FWPRec (Ours)0.7993 ±0.00210.2360 ±0.00200.2334 ±0.00220.2804 ±0.00233/3

KuaiRec (small_chronological)

Short-video watch logs, bounded small-matrix protocol; video captions provide the shared text vectors.

ModelAUCMRRnDCG@5nDCG@10seeds
Popularity0.14190.01950.00590.01011/1
ItemKNN0.36150.01890.00060.00151/1
BPRMF0.5567 ±0.00080.0670 ±0.00060.0444 ±0.00070.0645 ±0.00073/3
GRU4Rec0.5864 ±0.00440.0784 ±0.00340.0567 ±0.00440.0801 ±0.00493/3
SASRec0.5757 ±0.00930.0935 ±0.00370.0723 ±0.00400.0955 ±0.00483/3
DuoRec0.5294 ±0.00600.0628 ±0.00160.0407 ±0.00180.0583 ±0.00203/3
LRURec0.5926 ±0.00450.0821 ±0.00020.0611 ±0.00060.0860 ±0.00103/3
Mamba4Rec0.5824 ±0.00380.0765 ±0.00240.0550 ±0.00310.0796 ±0.00323/3
HSTU0.5811 ±0.01130.0785 ±0.00520.0571 ±0.00560.0801 ±0.00653/3
FWPRec (Ours)0.6106 ±0.00940.0889 ±0.00280.0678 ±0.00250.0941 ±0.00293/3

Where the learning rates came from

MIND-large and Amazon: the paper's stated values (default 10⁻³; 2·10⁻³ for GRU4Rec, DuoRec, LRURec, HSTU on MIND and for LRURec on Amazon). EB-NeRD, MovieLens-1M, KuaiRec: the paper names no settings, so every neural model ran a three-way screen (5·10⁻⁴, 10⁻³, 2·10⁻³) at seed 42, judged on validation nDCG@10 only — 72 screen runs whose winners are frozen in a checked-in selection file. Screens structurally cannot touch the test split.