Skip to main content
Skip to content
Gallery

Every architecture, side by side

All figures in one place, for comparing shapes at a glance. Hover a block to trace its connections; open a model for the full explanation.

Retrieval

Narrow millions of items down to a few thousand candidates.
Two-TowerMicrosoft · Google · YouTube · 2016
Open
offlineUser towerItem towerUser IDWatch historyContextItem IDAttributesContentEmbedding lookup + poolingmean-pool variable-length historyEmbedding lookupID ⊕ attribute ⊕ content featuresMLPReLU layers → dMLPReLU layers → dL2 normalizeL2 normalizeu ∈ ℝdv ∈ ℝd·ANN indexservingSampled softmaxin-batch negatives + log Q correction

Make retrieval a nearest-neighbor search by never letting user and item features meet until a single dot product.

MINDAlibaba · 2019
Open
iterateVMulti-interest extractor · B2I dynamic routingUser behaviorsitem embeddings e1, …, enShared bilinear mapûj|i = S eiDynamic routing (≈3 iterations)wij = softmaxj(bij); bij ← bij + ujT S eiSquash → interest capsulesuj = squash(Σi wij S ei)u1u2uKProfileConcat profile + ReLU layers→ K user vectors V = [v1, …, vK]Target itemeiLabel-aware attentionvu = V softmax(pow(VT ei, p))Sampled softmaxvu vs target itemServing: K ANN queriesunion / merge top-N per interest

Replace the single user vector of two-tower retrieval with several interest vectors, so a user who likes both hiking gear and baby products can retrieve both.

Feature Interaction

Rank candidates by learning how features combine.
Wide & DeepGoogle · 2016
Open
ywideydeepWide · memorizationDeep · generalizationCross-product transformsAND(installed=X, impressed=Y)Categorical featuresContinuous featuresLinear modelwT [x, φ(x)] + bEmbeddings32-d per categorical featureConcat~1200-dReLU · 1024ReLU · 512ReLU · 256+σP(install | x)

Train a memorizing linear model and a generalizing deep network jointly, so each covers the other's failure mode.

DeepFMHuawei · 2017
Open
raw x[e1,…,em]⋯FM componentDeep componentField 1Field 2Field jField mShared embedding layerei = Vi xi ∈ ℝk (one per field)Addition1st order: ⟨w, x⟩Inner products2nd order: Σi<j ⟨ei, ej⟩Hidden layer 1σ(W(1) a(0) + b(1))Hidden layer H+σŷ = σ(yFM + yDNN)

Replace hand-engineered crosses with a factorization machine that shares its embeddings with the deep network.

xDeepFMUSTC · Microsoft · 2018
Open
X0Xk+1CIN · Compressed Interaction NetworkSparse fieldsField embeddingsX0 ∈ ℝm × DHadamard with X⁰Zk+1h,i,* = Xkh,* ∘ X0i,*Compress (CNN-like filters)Xk+1h,* = Σi,j Wk,hi,j (Xki,* ∘ X0j,*)Sum-pool each layerpk = ΣD Xk → [p1, …, pT]Linearraw featuresDNNimplicit, bit-wise+σŷ

Generate explicit, bounded-degree feature interactions at the vector-wise level (whole embeddings interact), not the bit-wise level used by DCN-v1 and MLPs.

AutoIntPeking University · Mila · 2019
Open
WRes emInteracting layer× LField 1Field 2Field 3Field MEmbedding (one vector per field)categorical: Vm xm · numerical: vm · xme1e2e3eMMulti-head self-attention over fieldsα(h)m,k = softmaxk(⟨WQ em, WK ek⟩)+ReLUConcat all fieldsLinear → σŷ = σ(wT[eRes1; …; eResM] + b)

Use multi-head self-attention over feature-field embeddings to learn which high-order interactions matter, explicitly and interpretably.

DLRMMeta · 2019
Open
xdenseEmbedding tables · model-parallelDense featuresSparse 1Sparse 2Sparse KBottom MLPℝn_{dense} → ℝdTable 1lookup → dTable 2lookup → dTable Klookup → dFeature interaction⟨xi, xj⟩ for all pairs i < jConcat[xdense, dots]Top MLP→ 1σCTR

A deliberately simple, hardware-aware reference design: embeddings + bottom MLP, pairwise dot products, top MLP.

DCN-v2Google · 2020
Open
x0xl⋯xLhLCross networkexplicitDeep networkimplicitSparse & dense featuresEmbedding & stackingx0 = [e1; … ; ek; xdense] ∈ ℝdxlLinearWl xl + bl⊙+Cross layers 2 … Lsame form, x0 re-injectedDense layer 1ReLU(W h + b)Dense layer 2ReLU(W h + b)Dense layer LReLU(W h + b)ConcatLogit → sigmoidŷ = σ(wT[xL ; hL])

Learn bounded-degree feature crosses explicitly with a cheap cross layer, x₀ ⊙ (Wxₗ + b) + xₗ, instead of hoping an MLP finds them.

User Behavior Sequences

Model what the user did recently, conditioned on the candidate.
DINAlibaba · 2018
Open
vUvAAttention pooling over behaviorsInside an activation unitUser profile+ context featurese1e2eTvAcandidate adAct. unitw1Act. unitw2Act. unitwTWeighted sum poolingvU(A) = Σi wi · eiConcatMLP 200 → 80PReLU / Dice activationsσCTR[ ei ; ei ⊙ vA ; vA ]Small MLPDice · 36 unitswi ∈ ℝ

Let the candidate decide which parts of the user's history matter: attention-weighted pooling of behaviors conditioned on the target ad.

SASRecUC San Diego · 2018
Open
Self-attention block× bi1i2i3itItem embedding + learned positionÊk = Mi_k + PkCausal self-attentionsoftmax(QKT / √d + Mcausal) V+LayerNormPoint-wise FFNReLU(S W1 + b1) W2 + b2+LayerNormHidden statesF(b)1, …, F(b)tNext-item scoreri,k = F(b)k · Mi (target: ik+1)

Model the user's sequence with a causal (GPT-style) self-attention Transformer and train next-item prediction at every position.

DIENAlibaba · 2019
Open
⋯⋯h'TeaInterest extractorInterest evolvinge1e2e3eTeatarget adGRUGRUGRUGRUAux. lossht vs et+1, e-a1a2a3aTAUGRUAUGRUAUGRUAUGRUProfile + contextConcatMLP (Dice / PReLU)σCTR

Model interests as a hidden process that evolves over time, rather than as a bag of past behaviors, and let the candidate steer which evolution path matters.

BERT4RecAlibaba · 2019
Open
Transformer layer× Lv1[mask]v3v4[mask]Item + position embeddingh0i = vi + piBidirectional multi-head self-attentionevery position attends to every positionAdd & NormPosition-wise FFN (GELU)Add & NormhL2hL5Predict masked itemssoftmax(GELU(h WP + bP) ET + bO)

Model user sequences bidirectionally by training with a Cloze (masked-item) objective instead of left-to-right next-item prediction.

SIMAlibaba · 2020
Open
caea (query)Stage 1 · General Search Unit (cheap)Stage 2 · Exact Search Unit (precise)Lifelong behaviorsup to ~54k clicks / purchases over yearsCandidate itemcategory ca, embedding eaHard searchkeep behaviors with category = caSoft search (alt.)top-K by ⟨Wb ei, Wa ea⟩Top-K relevant behaviorsK ≈ a few hundredBehavior + time-gap embeddingzj = [ej ; eΔt_j]Multi-head target attentionAttention(q = ea, K = V = Z)Concat with short-term & other featuresMLP → CTR

Make target attention scale to lifelong histories (tens of thousands of behaviors) by first searching the history for candidate-relevant behaviors, then attending only to those.

Multi-Task

Predict several objectives (click, like, watch time) at once.
ESMMAlibaba · 2018
Open
Main task · CVRAuxiliary task · CTRUser featuresItem featuresShared embedding lookupCVR borrows representations learned from abundant click dataCVR towerMLPCTR towerMLPpCVRpCTR×pCTCVR = pCTR × pCVRLoss: CTCVRlabel: click & convert · all impressionsLoss: CTRlabel: click · all impressions

Fix sample selection bias in conversion-rate prediction by never training CVR on clicked samples alone: supervise pCTR and pCTCVR = pCTR × pCVR over the entire impression space.

MMoEGoogle · 2018
Open
gAgBShared expertsInput representation xExpert 1f1(x)Expert 2f2(x)Expert 3f3(x)Expert nfn(x)Gate Asoftmax(WA x)Gate Bsoftmax(WB x)ΣΣTower ATower BTask A · e.g. clickTask B · e.g. watch time

Share a pool of experts across tasks but give every task its own gate, so the model learns how much to share.

PLETencent · 2020
Open
Extraction layer 1 · CGCExtraction layers 2 … L× (L−1)Input embeddingsTask-A expertsEA,1, …, EA,m_AShared expertsES,1, …, ES,m_STask-B expertsEB,1, …, EB,m_BGate AGate SGate BhA(1)hS(1)hB(1)Same CGC structure, inputs from layer belowtask experts read hA / hB, shared experts read hS; the final layer drops the shared gateTower ATower BTask ATask B

Explicitly separate task-specific experts from shared experts, and refine the separation progressively over several layers.

Generative

Treat recommendation as generation: semantic IDs, LLM backbones, whole sessions and pages.
TIGERGoogle DeepMind · 2023
Open
tokenizecross-attn① Semantic IDs (offline)② Generative retrievalItem contenttitle, description, category, price…Sentence-T5 encoderx ∈ ℝ768RQ-VAE encoderz = E(x)Residual quantizationcd = argmink ‖rd − edk‖rd+1 = rd − edc_dSemantic ID(c1, c2, c3) + c4 for collisionsUser history as tokens(5,23,55) (5,25,78) … + user tokenTransformer encoderbidirectional over history tokensTransformer decoderautoregressive · cross-attends to encoderBeam search over codewordsc1 → c2 → c3Next itemdecoded semantic ID → item lookup

Turn retrieval into generation: give every item a semantic ID (a short code tuple) and have a seq2seq Transformer generate the next item's code.

HSTUMeta · 2024
Open
XHSTU layer× LΦ0a0Φ1a1ΦnanEmbedding layerone chronological sequence of length nPointwise projectionU, V, Q, K = Split(φ1(f1(X)))QKVUrabp,tposition + time biasSpatial aggregationφ2(QKT + rabp,t) VLayerNorm⊙Linearf2(·)+Next-token headsretrieval: next item Φi+1 · ranking: action ai at item positions

Recast recommendation as generative sequential transduction over raw user actions, with an attention block redesigned for recsys data, and show it follows scaling laws.

HLLMByteDance · 2024
Open
for every itemItem LLM · feature extractionUser LLM · interest modelingItem text + [ITEM]title · tags · descriptionItem LLMpretrained decoder, fully fine-tunedEihidden state at [ITEM]E1E2E3EnUser LLMcausal over item embeddings · word vocab removedÊ2Ê3Ê4Ên+1Generative lossInfoNCE(Êt+1, Et+1, negatives)Discriminative headfuse user seq + target item → BCE

Split the recommender into an Item LLM and a User LLM, both initialized from pretrained LLMs, so world knowledge helps item understanding while sequences stay short (one embedding per item).

OneRecKuaishou · 2025
Open
K, VvocabsamplesUser behavior pathwaysEncoder× L_encDecoder× L_decItem tokenizerRL post-trainingStaticuid · age · genderShort-termlast 20 viewsPositive feedback256 engaged viewsLifelong100k → clusters → 128 qConcat + positionalz(1) = [hu; hs; hp; hl] + eposDecoder input[BOS], s1, s2, s3 of target videoSelf-attention + FFNbidirectional over user tokensCausal self-attentionCross-attentionQ: decoder · K, V: encoderMoE FFNtop-k of N experts · loss-free balancingSoftmax over codebooknext semantic-ID token sj+1Generated videosbeam search: s¹ → s² → s³ per videoMultimodal LLM + QFormercaptions, frames, ASR/OCR → 4 tokensRQ-KMeans · 3 × 8192coarse-to-fine semantic IDsReward systemP-Score preference + format + businessECPOearly-clipped GRPO on sampled outputs

Collapse the multi-stage recommendation cascade into a single generative model that directly decodes which videos to show, and make it cheaper to run than the cascade it replaces.

MTGRMeituan · 2025
Open
One sample = one user + K candidatesMTGR block (HSTU-style)× LU1UmS1SnR1Rr[C,I]1[C,I]KFeature → token embeddingeach feature / item → dmodel via MLPGroup LayerNormseparate statistics per token group: U · S · R · CProjectionQ, K, V, U = SiLU(f1(X))Masked SiLU attentionSiLU(QKT) V with dynamic maskDynamic maskU, S: visible to allR: causal · C: self onlyGate + outputf2(GLN(AV) ⊙ U) + XCandidate token statesone final vector per candidatePer-candidate headsCTR · CTCVR for all K candidates in one pass

Get generative-recommender scaling without discarding the hand-crafted cross features that DLRMs depend on: put the candidates and their cross features into the token sequence.

GenRec (JD)JD.com · 2026
Open
samples① Semantic IDsDecoder-only LLM · Qwen2.5-3B× 36② TrainingItem image + textQwen2.5-VL encodermultimodal item embeddingRQ-KMeansSID(v) = (s1, s2, s3)User history SuSID(v1) SID(v2) … SID(vn) — 3 tokens per itemAsymmetric Token Mergerhv = Linear([e(s1); e(s2); e(s3)]) — prompt side onlyCausal Transformerdeep & narrow: 36 layers × 2048 hiddenGenerate a page of SIDsYpage = [SID(v) : v ∈ O ∪ C ∪ E], sorted by intensityBeam searchbeam-width items per requestSFT: page-wise NTP−Σt log P(yt | Su, y<t)GRPO-SRgroup-relative RL + NLL anchorPreference rewardSIM-based scorer · gated by SID validity

Make semantic-ID generative retrieval practical at e-commerce scale with three fixes: supervise pages instead of single items, compress the prompt, and align with a stabilized RL objective.

GenRec (Netflix)Netflix · 2026
Open
trainDecoder-only LLM · 1B–10B× LCatalog-aware ranking headWatch history HCandidate metadata {Mᵢ}Task instruction τVerbalizer V(·)keep high-signal events · compress binges · drop low-signal views~5,000 → ~1,700 tokens per promptText prompt xx = V(H, {Mi}i ∈ C, τ)Pretrained open LLMPhase 1: adapt to Netflix data · Phase 2: ranking post-trainingh ∈ ℝdpooled hidden state·Item embeddingsei ∈ ℝd for every catalog itemSoftmax over the catalogp(i | x) = softmaxi(h · eiT)Training lossL = αLrank + βLLM + γLmiscRank all candidates in one passvLLM prefill-only serving · no decoding loop

Use an open LLM as the ranking backbone, but replace text generation with a catalog-aware scoring head, so one prefill pass ranks every candidate.

GenPageNetflix · 2026
Open
appendorPrompt · member & request contextResponse · the generated pageDecoder-only Transformer× LTraining recipeProfileRequesth1h2hTR1e1,1e1,2R2e2,1e2,2Domain-specific token embeddingone token per entity, row, action type & context valueCausal self-attention + FFNprompt → page, left-to-right, top-to-bottomOutput head (untied)logits over the row + entity vocabularyConstrained decodingmask: business rules · deduprow grammarHybrid row decodingfirst k entities step-by-steprest of the row in one passNext tokena row token or an entity tokenPretrainingnext-token prediction onengaged production pagesPost-train: WBCreward-signed, weightedbinary target per tokenPost-train: RLDr. GRPO vs. page rewardmodel + KL penalty

Replace the multi-stage page-construction stack with a single model that generates the entire structured homepage autoregressively, trained like an LLM: pretraining, then post-training.

Scaling Backbones

Transformer-like ranking backbones designed to scale with compute.
WukongMeta · 2024
Open
Xi (projected if shapes differ)Xi+1Wukong layer× LDense & sparse featuresEmbedding layerX0 ∈ ℝn × dXiFactorization Machine BlockMLP(LN(flatten(Xi XiT))) → nF × dLinear Compress BlockWL Xi → nL × dConcat(nF + nL) × d+LayerNormMLP head → predictionsflatten(XL)

Stack factorization-machine blocks so interaction order grows exponentially with depth, giving recommendation its first clean scaling law.

RankMixerByteDance · 2025
Open
residualRankMixer block× LHeterogeneous featuresuser · item · context · sequence summariesSemantic grouping → tokenizationxi = Proj(concat(groupi)) ∈ ℝD, i = 1…Tx1x2x3xTMulti-head token mixingsplit each token into H heads → head h of all tokens forms new token h+LayerNormFFN1FFN2FFN3FFNT+LayerNormMean pooling → task towers

Design the ranking backbone for GPUs: parameter-free token mixing plus per-token FFNs, scaling to billion-parameter models within the same latency.

OneTransByteDance · NTU · 2025
Open
OneTrans block× LSequential featuresbehavior sequences: clicks, purchases, …Non-sequential featuresuser profile, candidate item, contextSequential tokenizerper-event embedding · sequences joined by [SEP]Non-sequential tokenizergroup-wise or auto-split → k tokensS1S2SEPSmNS1NS2NSkRMSNormMixed causal attentionS-tokens: shared WQ,K,VNS-token i: own W(i)Q,K,VKV cacheS-token K/V+RMSNormMixed FFNshared FFN for S-tokens · token-specific FFN per NS-token+Pyramid truncationkeep queries for the most recent S-tokens + all NS-tokensFinal NS-token statesonly NS-tokens remain after the last layerTask towersCTR · CVR · …

Replace the separate 'sequence module' and 'feature-interaction module' with one causal Transformer over a unified token sequence.