FWPRec Experiments
All experiments
Efficiency benchmark · one RTX 4080 SUPER · updated 2026-09-24 11:24 UTC

What each model costs

The same measurements efficiency-focused recommender papers publish — trainable parameters, peak GPU memory, training time per epoch, and inference time per batch — taken on one idle GPU with each model in its exact published main-table configuration. Timings are medians over repeated, synchronized steps; the last column is the real wall-clock of the published three-seed runs.

Takeaway

The notepad's inference cost sits between the light recurrent models and the attention stack, with the smallest inference memory of any adaptive model — its per-user state is bounded, so serving memory stays flat while attention models grow with history. Training is mid-pack: the write loop makes epochs slower than GRU4Rec's but comparable to the state-space and contrastive baselines. The fastest trainer everywhere is GRU4Rec; the slowest is Mamba4Rec's unfused scan.

MIND-large · batch 1024

ModelTrainable paramsTrain s/epoch Peak train GBInfer ms/batchPeak infer GB Total train min†
Popularity0··0.040.0043.7
ItemKNN0··942.390.00315.7
BPRMF48.03M13.060.9210.110.40912.1
GRU4Rec25K8.960.4472.320.07810.6
SASRec103K21.720.6214.180.18111.2
DuoRec103K49.711.6533.970.69314.9
LRURec101K87.051.2654.720.24222.9
Mamba4Rec66K100.541.7186.650.28218.7
HSTU62K28.830.6633.540.2739.5
FWPRec (Ours)83K50.411.50910.860.19116.1

EB-NeRD · batch 1024

ModelTrainable paramsTrain s/epoch Peak train GBInfer ms/batchPeak infer GB Total train min†
Popularity0··0.050.0030.9
ItemKNN0··1015.770.0034.4
BPRMF1.20M0.270.0710.090.0391.5
GRU4Rec25K1.060.081.760.0741.3
SASRec103K2.280.5994.050.1591.4
DuoRec103K5.21.6313.910.6711.2
LRURec101K8.91.2434.560.221.7
Mamba4Rec66K10.091.6966.480.261.7
HSTU62K2.860.6423.410.2511.3
FWPRec (Ours)83K5.151.48710.780.1692.2

Amazon Video Games · batch 256

ModelTrainable paramsTrain s/epoch Peak train GBInfer ms/batchPeak infer GB Total train min†
Popularity0··0.030.0020.6
ItemKNN0··431.620.0021.2
BPRMF6.06M1.180.1210.080.0651.2
GRU4Rec25K3.940.0921.590.0281.3
SASRec106K6.080.3322.330.1241.5
DuoRec106K14.520.8742.210.3982.6
LRURec101K24.240.6534.980.1234.1
Mamba4Rec66K32.520.8598.360.1386.1
HSTU65K10.970.412.410.1732.3
FWPRec (Ours)83K14.480.3993.340.0692.6

MovieLens-1M · batch 256

ModelTrainable paramsTrain s/epoch Peak train GBInfer ms/batchPeak infer GB Total train min†
Popularity0··0.030.0011.2
ItemKNN0··524.480.0014.4
BPRMF387K2.030.0530.070.0172.9
GRU4Rec25K8.710.0532.580.0243.2
SASRec103K8.850.1641.480.0534.0
DuoRec103K21.580.421.380.1844.7
LRURec101K31.110.3283.570.0697.0
Mamba4Rec66K44.250.4384.170.0788.2
HSTU62K10.530.1741.550.0763.5
FWPRec (Ours)83K32.030.393.490.067.2

KuaiRand-Pure (IDs) · batch 512

ModelTrainable paramsTrain s/epoch Peak train GBInfer ms/batchPeak infer GB Total train min†
Popularity0··0.040.0032.5
ItemKNN0··567.190.0032.8
BPRMF2.22M1.310.0530.120.033.9
GRU4Rec516K5.030.0692.40.0474.9
SASRec597K11.990.644.060.2215.5
DuoRec597K27.091.7214.060.7656.2
LRURec592K71.681.2795.680.2219.4
Mamba4Rec557K68.731.6997.690.24920.0
HSTU556K20.570.7984.640.328.6
FWPRec (Ours)567K13.130.7633.350.1026.0

KuaiRec · batch 256

ModelTrainable paramsTrain s/epoch Peak train GBInfer ms/batchPeak infer GB Total train min†
Popularity0··0.040.0024.8
ItemKNN0··619.250.00227.1
BPRMF90K5.230.0490.130.0156.5
GRU4Rec25K27.080.0572.890.0297.8
SASRec106K30.890.3262.540.1199.0
DuoRec106K73.370.8692.270.39314.0
LRURec101K142.00.6484.950.11826.0
Mamba4Rec66K248.290.8548.330.13235.7
HSTU65K46.40.4052.330.1689.3
FWPRec (Ours)83K96.520.3933.30.06412.4

Fine print

† mean wall-clock of the three published final seeds, including per-epoch validation and early stopping — so it reflects what the runs actually cost, not epochs × epoch-time. Trainable parameters exclude the frozen Qwen table (identical for every model on content columns); on KuaiRand all item tables are learned, so counts include them. Popularity and ItemKNN involve no gradient training; their fit happens at construction. Batch sizes are each column's published training batch size.