Click-through prediction for a news recommendation task. Each user is shown a slate of candidate news articles; the model scores every article in the slate and predicts which ones the user will click. Three architectures are implemented on a shared BERT-embedding pipeline so they can be compared under identical data and evaluation.
DM_final_v2-master/holds the original 2024 submission; the top-level scripts are a cleaned-up 2026 rewrite of the same pipeline.
Embeddings. News text from train_news.tsv is encoded with BERT into 768-d article
embeddings. A user embedding is the mean of the embeddings of the articles that user has
clicked, taken from train_behaviors.tsv — so a user is represented in the same semantic
space as the items.
Task framing. For each impression, the model receives one user embedding and a
variable-length stack of candidate news embeddings, and emits a click probability per
candidate. Trained with BCELoss, evaluated with ROC-AUC on a 9:1 train/validation split.
| Script | Model | Architecture |
|---|---|---|
ncf.py |
NCF | Separate user/news MLP towers (768 → 128, ReLU, dropout 0.3) → concat → 128 → 64 → 1, sigmoid |
ncf_layer.py |
NCF-Layer | Deeper towers (768 → 128 → 64, dropout 0.1) → concat → MLP → 1, sigmoid |
ncf_bert4rec.py |
BERT4Rec | User embedding prepended to the candidate stack, passed through a 2-layer TransformerEncoder (8 heads, FFN 4×, dropout 0.2); per-candidate logits from the news positions |
The NCF variants score each candidate independently. BERT4Rec lets self-attention run across the whole slate, so candidates are scored in context of each other rather than in isolation — that is the substantive difference between the two families here.
| NCF | NCF-Layer | BERT4Rec | |
|---|---|---|---|
| Optimizer | Adam, lr 1e-3, wd 1e-5 | Adam | AdamW, lr 1e-3, wd 1e-5 |
| Scheduler | ReduceLROnPlateau (max AUC, ×0.5, patience 2) |
step | CosineAnnealingWarmRestarts (T₀=10, T_mult=2) |
| Batch size | 32 train / 128 val | 32 / 128 | 32 / 128 |
| Epochs | 100, best-AUC checkpointing | 100 | 100 |
Runs on Apple Silicon (mps) by default; switch the device line for CUDA.
From the written report (report.pdf) — Kaggle public leaderboard score, comparing model
family against loss function:
| NCF | GRU4Rec | |
|---|---|---|
| BPR | 66.79 | 63.25 |
| BCE | 65.14 | 65.33 |
NCF pairs best with BPR: a ranking loss suits an architecture built to compare many items for one user. The sequential model gains nothing from BPR, since it is built to predict the next item rather than to rank a slate — with BCE it edges ahead instead.
ncf.py # NCF — train + eval + plots
ncf_layer.py # NCF-Layer variant
ncf_bert4rec.py # BERT4Rec (Transformer over the candidate slate)
submission.py # Loads best checkpoints, writes predictions for submission
DM_final_v2-master/ # Original 2024 submission (inference scripts)
report.pdf # Written report: method, ablation, analysis
data/ # BERT embeddings + behaviour TSVs (gitignored, ~1.9 GB)
model/ # Saved checkpoints (gitignored, ~1.8 GB)
pip install torch pandas scikit-learn matplotlib tqdm
python ncf.py # or ncf_layer.py / ncf_bert4rec.py
python submission.py # writes predictions from the best checkpointsdata/ and model/ are not tracked — the embedding pickles and checkpoints are several
gigabytes. Generate the embeddings with BERT over train_news.tsv before training.