What each model costs
The same measurements efficiency-focused recommender papers publish — trainable parameters, peak GPU memory, training time per epoch, and inference time per batch — taken on one idle GPU with each model in its exact published main-table configuration. Timings are medians over repeated, synchronized steps; the last column is the real wall-clock of the published three-seed runs.
The notepad's inference cost sits between the light recurrent models and the attention stack, with the smallest inference memory of any adaptive model — its per-user state is bounded, so serving memory stays flat while attention models grow with history. Training is mid-pack: the write loop makes epochs slower than GRU4Rec's but comparable to the state-space and contrastive baselines. The fastest trainer everywhere is GRU4Rec; the slowest is Mamba4Rec's unfused scan.
MIND-large · batch 1024
| Model | Trainable params | Train s/epoch | Peak train GB | Infer ms/batch | Peak infer GB | Total train min† |
|---|---|---|---|---|---|---|
| Popularity | 0 | · | · | 0.04 | 0.004 | 3.7 |
| ItemKNN | 0 | · | · | 942.39 | 0.003 | 15.7 |
| BPRMF | 48.03M | 13.06 | 0.921 | 0.11 | 0.409 | 12.1 |
| GRU4Rec | 25K | 8.96 | 0.447 | 2.32 | 0.078 | 10.6 |
| SASRec | 103K | 21.72 | 0.621 | 4.18 | 0.181 | 11.2 |
| DuoRec | 103K | 49.71 | 1.653 | 3.97 | 0.693 | 14.9 |
| LRURec | 101K | 87.05 | 1.265 | 4.72 | 0.242 | 22.9 |
| Mamba4Rec | 66K | 100.54 | 1.718 | 6.65 | 0.282 | 18.7 |
| HSTU | 62K | 28.83 | 0.663 | 3.54 | 0.273 | 9.5 |
| FWPRec (Ours) | 83K | 50.41 | 1.509 | 10.86 | 0.191 | 16.1 |
EB-NeRD · batch 1024
| Model | Trainable params | Train s/epoch | Peak train GB | Infer ms/batch | Peak infer GB | Total train min† |
|---|---|---|---|---|---|---|
| Popularity | 0 | · | · | 0.05 | 0.003 | 0.9 |
| ItemKNN | 0 | · | · | 1015.77 | 0.003 | 4.4 |
| BPRMF | 1.20M | 0.27 | 0.071 | 0.09 | 0.039 | 1.5 |
| GRU4Rec | 25K | 1.06 | 0.08 | 1.76 | 0.074 | 1.3 |
| SASRec | 103K | 2.28 | 0.599 | 4.05 | 0.159 | 1.4 |
| DuoRec | 103K | 5.2 | 1.631 | 3.91 | 0.671 | 1.2 |
| LRURec | 101K | 8.9 | 1.243 | 4.56 | 0.22 | 1.7 |
| Mamba4Rec | 66K | 10.09 | 1.696 | 6.48 | 0.26 | 1.7 |
| HSTU | 62K | 2.86 | 0.642 | 3.41 | 0.251 | 1.3 |
| FWPRec (Ours) | 83K | 5.15 | 1.487 | 10.78 | 0.169 | 2.2 |
Amazon Video Games · batch 256
| Model | Trainable params | Train s/epoch | Peak train GB | Infer ms/batch | Peak infer GB | Total train min† |
|---|---|---|---|---|---|---|
| Popularity | 0 | · | · | 0.03 | 0.002 | 0.6 |
| ItemKNN | 0 | · | · | 431.62 | 0.002 | 1.2 |
| BPRMF | 6.06M | 1.18 | 0.121 | 0.08 | 0.065 | 1.2 |
| GRU4Rec | 25K | 3.94 | 0.092 | 1.59 | 0.028 | 1.3 |
| SASRec | 106K | 6.08 | 0.332 | 2.33 | 0.124 | 1.5 |
| DuoRec | 106K | 14.52 | 0.874 | 2.21 | 0.398 | 2.6 |
| LRURec | 101K | 24.24 | 0.653 | 4.98 | 0.123 | 4.1 |
| Mamba4Rec | 66K | 32.52 | 0.859 | 8.36 | 0.138 | 6.1 |
| HSTU | 65K | 10.97 | 0.41 | 2.41 | 0.173 | 2.3 |
| FWPRec (Ours) | 83K | 14.48 | 0.399 | 3.34 | 0.069 | 2.6 |
MovieLens-1M · batch 256
| Model | Trainable params | Train s/epoch | Peak train GB | Infer ms/batch | Peak infer GB | Total train min† |
|---|---|---|---|---|---|---|
| Popularity | 0 | · | · | 0.03 | 0.001 | 1.2 |
| ItemKNN | 0 | · | · | 524.48 | 0.001 | 4.4 |
| BPRMF | 387K | 2.03 | 0.053 | 0.07 | 0.017 | 2.9 |
| GRU4Rec | 25K | 8.71 | 0.053 | 2.58 | 0.024 | 3.2 |
| SASRec | 103K | 8.85 | 0.164 | 1.48 | 0.053 | 4.0 |
| DuoRec | 103K | 21.58 | 0.42 | 1.38 | 0.184 | 4.7 |
| LRURec | 101K | 31.11 | 0.328 | 3.57 | 0.069 | 7.0 |
| Mamba4Rec | 66K | 44.25 | 0.438 | 4.17 | 0.078 | 8.2 |
| HSTU | 62K | 10.53 | 0.174 | 1.55 | 0.076 | 3.5 |
| FWPRec (Ours) | 83K | 32.03 | 0.39 | 3.49 | 0.06 | 7.2 |
KuaiRand-Pure (IDs) · batch 512
| Model | Trainable params | Train s/epoch | Peak train GB | Infer ms/batch | Peak infer GB | Total train min† |
|---|---|---|---|---|---|---|
| Popularity | 0 | · | · | 0.04 | 0.003 | 2.5 |
| ItemKNN | 0 | · | · | 567.19 | 0.003 | 2.8 |
| BPRMF | 2.22M | 1.31 | 0.053 | 0.12 | 0.03 | 3.9 |
| GRU4Rec | 516K | 5.03 | 0.069 | 2.4 | 0.047 | 4.9 |
| SASRec | 597K | 11.99 | 0.64 | 4.06 | 0.221 | 5.5 |
| DuoRec | 597K | 27.09 | 1.721 | 4.06 | 0.765 | 6.2 |
| LRURec | 592K | 71.68 | 1.279 | 5.68 | 0.22 | 19.4 |
| Mamba4Rec | 557K | 68.73 | 1.699 | 7.69 | 0.249 | 20.0 |
| HSTU | 556K | 20.57 | 0.798 | 4.64 | 0.32 | 8.6 |
| FWPRec (Ours) | 567K | 13.13 | 0.763 | 3.35 | 0.102 | 6.0 |
KuaiRec · batch 256
| Model | Trainable params | Train s/epoch | Peak train GB | Infer ms/batch | Peak infer GB | Total train min† |
|---|---|---|---|---|---|---|
| Popularity | 0 | · | · | 0.04 | 0.002 | 4.8 |
| ItemKNN | 0 | · | · | 619.25 | 0.002 | 27.1 |
| BPRMF | 90K | 5.23 | 0.049 | 0.13 | 0.015 | 6.5 |
| GRU4Rec | 25K | 27.08 | 0.057 | 2.89 | 0.029 | 7.8 |
| SASRec | 106K | 30.89 | 0.326 | 2.54 | 0.119 | 9.0 |
| DuoRec | 106K | 73.37 | 0.869 | 2.27 | 0.393 | 14.0 |
| LRURec | 101K | 142.0 | 0.648 | 4.95 | 0.118 | 26.0 |
| Mamba4Rec | 66K | 248.29 | 0.854 | 8.33 | 0.132 | 35.7 |
| HSTU | 65K | 46.4 | 0.405 | 2.33 | 0.168 | 9.3 |
| FWPRec (Ours) | 83K | 96.52 | 0.393 | 3.3 | 0.064 | 12.4 |
Fine print
† mean wall-clock of the three published final seeds, including per-epoch validation and early stopping — so it reflects what the runs actually cost, not epochs × epoch-time. Trainable parameters exclude the frozen Qwen table (identical for every model on content columns); on KuaiRand all item tables are learned, so counts include them. Popularity and ItemKNN involve no gradient training; their fit happens at construction. Batch sizes are each column's published training batch size.