Recommender system architectures
27 influential models, from two-tower retrieval to generative, LLM-style recommenders that write whole pages. Each one is drawn as an interactive figure: click any block for details, or step through the forward pass.
Where each model sits
Production recommenders are a funnel. Each stage trades model cost against the number of items it scores.
Timeline
Roughly: hand-crafted crosses → learned interactions → sequences & attention → scalable Transformer backbones → generative, LLM-style recommenders.
Retrieval
Narrow millions of items down to a few thousand candidates.Encode users and items separately, then score by dot product, so retrieval becomes nearest-neighbor search.
View architectureRepresent each user with K interest vectors, extracted from behaviors by capsule-style dynamic routing, and retrieve with each one.
View architectureFeature Interaction
Rank candidates by learning how features combine.A linear model over hand-crafted crosses for memorization, plus an MLP over embeddings for generalization, trained jointly.
View architectureReplace Wide & Deep's hand-made crosses with a Factorization Machine that shares embeddings with the deep tower.
View architectureA Compressed Interaction Network learns explicit high-order crosses at the vector level, alongside a linear part and a DNN.
View architectureTreat each feature field as a token and let multi-head self-attention decide which fields should interact.
View architectureDense features through a bottom MLP, sparse features through embedding tables, all pairwise dot products, then a top MLP.
View architectureExplicit, bounded-degree feature crosses learned by a stack of cheap cross layers.
View architectureUser Behavior Sequences
Model what the user did recently, conditioned on the candidate.Pool the user's behavior history with attention weights computed against the candidate ad, so the user vector depends on what is being scored.
View architectureA decoder-only (GPT-style) Transformer over the user's item sequence, trained to predict the next item at every position.
View architectureExtract a latent interest at every step with a GRU, then let the interests relevant to the candidate evolve through an attention-gated GRU.
View architectureBERT for item sequences: mask random items and predict them using context from both directions.
View architectureTwo-stage attention over lifelong histories: a cheap search keeps only the behaviors related to the candidate, then precise attention runs on those.
View architectureMulti-Task
Predict several objectives (click, like, watch time) at once.Learn conversion rate indirectly through pCTR × pCVR, so both losses are defined over all impressions.
View architectureShared experts, with one softmax gate per task deciding how to mix them.
View architectureSeparate task-specific and shared experts explicitly, and stack the extraction across layers.
View architectureGenerative
Treat recommendation as generation: semantic IDs, LLM backbones, whole sessions and pages.Give every item a short tuple of semantic codewords, then have a seq2seq Transformer generate the next item's code directly.
View architectureReformulate retrieval and ranking as sequential transduction over (item, action) tokens, with an attention block redesigned for recommendation data.
View architectureTwo stacked LLMs: one reads an item's text and outputs its embedding, the other reads a sequence of those embeddings and predicts the next one.
View architectureOne encoder–decoder with MoE replaces the whole retrieve → pre-rank → rank cascade, generating semantic IDs of videos and aligned with RL.
View architectureAn HSTU-style Transformer ranker that keeps the old DLRM cross features, and scores all of a user's candidates in a single sequence.
View architectureA decoder-only LLM over semantic IDs that is trained to generate a whole page of interactions, then aligned with RL.
View architectureVerbalize the member's history into text, run a pretrained LLM once, and score the entire catalog with a dot-product head. No token-by-token decoding.
View architectureOne decoder-only Transformer reads the user's context as a prompt and writes the whole multi-row homepage, rows and titles, token by token.
View architectureScaling Backbones
Transformer-like ranking backbones designed to scale with compute.Stack factorization-machine blocks so that interaction order grows exponentially with depth, and scale it like an LLM.
View architectureA hardware-friendly Transformer variant: parameter-free token mixing instead of self-attention, and a separate FFN for every feature token.
View architecturePut behavior sequences and non-sequential features into one token sequence and run a single causal Transformer with mixed parameterization.
View architecture