FWPRec Experiments
All experiments
MovieLens-1M · Amazon · KuaiRand-1K · seeds 42, 43 · four RTX cards

How does the notepad compare to the field?

Every model in this project is our own rewrite of someone else's, so a comparison is only worth anything if all of them get the same deal. We gave nineteen models one identical training budget, ran each twice with different random starts, and scored them the same way on three very different datasets: movie ratings, Amazon game purchases, and short-video watch logs. Nothing was tuned for our own method. The point was not to win — the notepad's case rests on being cheap and inspectable — but to find out honestly where it sits among models people actually publish.

Takeaway

It depends on the dataset, and we are reporting both halves. On MovieLens-1M the notepad is genuinely strong at finding the right film — 2nd of 17 on hit rate at 0.8269, behind only HSTU — though it orders its top ten less sharply, landing 9th on nDCG@10. On Amazon it drops to 11th of 17. The efficiency claim is the clean part: the bounded rank-16 notepad matches the full one on every dataset, so the smaller memory costs nothing.

MovieLens-1M

Ranked by nDCG@10 on the held-out test set. All runs finished.

HSTU
0.6419
Mamba4Rec
0.6336
LRURec
0.6310
SASRec
0.6259
BSARec
0.6235
TiSASRec
0.6224
DuoRec
0.6211
GRU4Rec
0.6179
#ModelnDCG@10HR@10
1HSTU0.64190.8359
2Mamba4Rec0.63360.8240
3LRURec0.63100.8247
4SASRec0.62590.8219
5BSARec0.62350.8212
6TiSASRec0.62240.8215
7DuoRec0.62110.8198
8GRU4Rec0.61790.8161
9FWPRec (dense)0.61480.8269
10FWPRec (rank-16)0.61200.8248
11LinRec0.58590.7982
12FPMC0.54700.7797
13SeqRules0.54000.7439
14BERT4Rec0.49180.6858
15BPRMF0.39130.6682
16ItemKNN0.32820.5780
17Popularity0.23860.4363
NRMSn/an/a
LSTURn/an/a

Amazon Video Games

Ranked by nDCG@10 on the held-out test set. All runs finished.

BSARec
0.1412
BERT4Rec
0.1411
SASRec
0.1383
TiSASRec
0.1379
GRU4Rec
0.1364
HSTU
0.1362
LRURec
0.1353
Mamba4Rec
0.1329
#ModelnDCG@10HR@10
1BSARec0.14120.2159
2BERT4Rec0.14110.2191
3SASRec0.13830.2162
4TiSASRec0.13790.2157
5GRU4Rec0.13640.2092
6HSTU0.13620.2115
7LRURec0.13530.2093
8Mamba4Rec0.13290.2070
9FPMC0.11230.2044
10SeqRules0.10970.1801
11FWPRec (dense)0.10610.1699
12FWPRec (rank-16)0.10560.1698
13DuoRec0.09890.1668
14ItemKNN0.09530.1607
15BPRMF0.08750.1469
16Popularity0.07450.1354
17LinRec0.07440.1352
NRMSn/an/a
LSTURn/an/a

KuaiRand-1K

Ranked by nDCG@10 on the held-out test set. 13 of 28 runs finished so far.

BPRMF
0.4077
Mamba4Rec
0.3631
GRU4Rec
0.3569
LinRec
0.3547
FWPRec (rank-16)
0.3531
FWPRec (dense)
0.3527
HSTU
0.3517
SeqRules
0.0914
#ModelnDCG@10HR@10
1BPRMF0.40770.8329
2Mamba4Rec0.36310.7613
3GRU4Rec0.35690.7523
4LinRec0.35470.6993
5FWPRec (rank-16)0.35310.7231
6FWPRec (dense)0.35270.7218
7HSTU0.35170.7505
8SeqRules0.09140.2004
9Popularity0.04260.0931
ItemKNNn/an/a
FPMCn/an/a
NRMSn/an/a
LSTURn/an/a

How it was run

75 of 90 runs finished so far. Each one records the code version, a checksum of the data it read, the exact settings, the GPU it used and how long it took, and each re-checks its own scores from its saved predictions before being counted.

Two seeds per model, reported as the mean; we do not quote a standard deviation from two numbers. The three datasets are scored against different-sized candidate lists, so a column is only comparable down its own length, never across. Rows marked n/a could not be run and say why in the record: NRMS and LSTUR need item text none of these datasets carry, ItemKNN's memory grows with the square of a user's history, and FPMC's three embedding tables over 4.4 million videos exceed the cards we have. KuaiRand-1K is still running.