Every architecture, side by side
All figures in one place, for comparing shapes at a glance. Hover a block to trace its connections; open a model for the full explanation.
Retrieval
Narrow millions of items down to a few thousand candidates.Make retrieval a nearest-neighbor search by never letting user and item features meet until a single dot product.
Replace the single user vector of two-tower retrieval with several interest vectors, so a user who likes both hiking gear and baby products can retrieve both.
Feature Interaction
Rank candidates by learning how features combine.Train a memorizing linear model and a generalizing deep network jointly, so each covers the other's failure mode.
Replace hand-engineered crosses with a factorization machine that shares its embeddings with the deep network.
Generate explicit, bounded-degree feature interactions at the vector-wise level (whole embeddings interact), not the bit-wise level used by DCN-v1 and MLPs.
Use multi-head self-attention over feature-field embeddings to learn which high-order interactions matter, explicitly and interpretably.
A deliberately simple, hardware-aware reference design: embeddings + bottom MLP, pairwise dot products, top MLP.
Learn bounded-degree feature crosses explicitly with a cheap cross layer, x₀ ⊙ (Wxₗ + b) + xₗ, instead of hoping an MLP finds them.
User Behavior Sequences
Model what the user did recently, conditioned on the candidate.Let the candidate decide which parts of the user's history matter: attention-weighted pooling of behaviors conditioned on the target ad.
Model the user's sequence with a causal (GPT-style) self-attention Transformer and train next-item prediction at every position.
Model interests as a hidden process that evolves over time, rather than as a bag of past behaviors, and let the candidate steer which evolution path matters.
Model user sequences bidirectionally by training with a Cloze (masked-item) objective instead of left-to-right next-item prediction.
Make target attention scale to lifelong histories (tens of thousands of behaviors) by first searching the history for candidate-relevant behaviors, then attending only to those.
Multi-Task
Predict several objectives (click, like, watch time) at once.Fix sample selection bias in conversion-rate prediction by never training CVR on clicked samples alone: supervise pCTR and pCTCVR = pCTR × pCVR over the entire impression space.
Share a pool of experts across tasks but give every task its own gate, so the model learns how much to share.
Explicitly separate task-specific experts from shared experts, and refine the separation progressively over several layers.
Generative
Treat recommendation as generation: semantic IDs, LLM backbones, whole sessions and pages.Turn retrieval into generation: give every item a semantic ID (a short code tuple) and have a seq2seq Transformer generate the next item's code.
Recast recommendation as generative sequential transduction over raw user actions, with an attention block redesigned for recsys data, and show it follows scaling laws.
Split the recommender into an Item LLM and a User LLM, both initialized from pretrained LLMs, so world knowledge helps item understanding while sequences stay short (one embedding per item).
Collapse the multi-stage recommendation cascade into a single generative model that directly decodes which videos to show, and make it cheaper to run than the cascade it replaces.
Get generative-recommender scaling without discarding the hand-crafted cross features that DLRMs depend on: put the candidates and their cross features into the token sequence.
Make semantic-ID generative retrieval practical at e-commerce scale with three fixes: supervise pages instead of single items, compress the prompt, and align with a stabilized RL objective.
Use an open LLM as the ranking backbone, but replace text generation with a catalog-aware scoring head, so one prefill pass ranks every candidate.
Replace the multi-stage page-construction stack with a single model that generates the entire structured homepage autoregressively, trained like an LLM: pretraining, then post-training.
Scaling Backbones
Transformer-like ranking backbones designed to scale with compute.Stack factorization-machine blocks so interaction order grows exponentially with depth, giving recommendation its first clean scaling law.
Design the ranking backbone for GPUs: parameter-free token mixing plus per-token FFNs, scaling to billion-parameter models within the same latency.
Replace the separate 'sequence module' and 'feature-interaction module' with one causal Transformer over a unified token sequence.