Two-TowerDual-Encoder Retrieval (DSSM / YouTube DNN)
Encode users and items separately, then score by dot product, so retrieval becomes nearest-neighbor search.
Sampling-Bias-Corrected Neural Modeling for Large Corpus Item Recommendations§1The key contribution
Make retrieval a nearest-neighbor search by never letting user and item features meet until a single dot product.
Scoring every item in a 10⁸-item corpus with a model that looks at user and item together is far too slow to do per request. Classic collaborative filtering (matrix factorization) is fast but cannot use rich features or handle new items.
Factor the model into two independent encoders. Item vectors can then be precomputed and indexed, and serving reduces to one user-tower pass plus an approximate max-inner-product search. Train it as a softmax over the whole corpus, approximated with in-batch negatives that are corrected for sampling bias.
Figure 1. Two-Tower architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
A user tower and an item tower each turn their features into a d-dimensional vector. The score is just ⟨u, v⟩. Because the item tower never sees the user, every item vector can be computed offline and loaded into an approximate nearest-neighbor (ANN) index. At request time you run the user tower once and fetch the top-k items in milliseconds from a corpus of hundreds of millions. Training treats it as extreme classification: a softmax over items, approximated with in-batch or sampled negatives and a log-Q correction for popularity bias.
Lineage. Origins in DSSM (Microsoft, 2013) for web search. YouTube's deep candidate generator (2016) and the 2019 sampling-bias-corrected paper made it the default retrieval architecture. Most production retrieval stacks still use a variant of it today.
§4Key equations
Q(y) is the probability that item y is sampled into a batch (roughly its popularity).
§5Why it works
- Decoupling makes serving cost independent of corpus size: item vectors are precomputed, and the user side runs once.
- The dot-product bottleneck is the price: no feature crosses between user and item, so the model can only express 'similarity in a d-dim space'.
- This is why recommendation is a funnel: two-tower retrieval → heavier ranking models (DCN, DIN, RankMixer…) on a few thousand candidates.
§6Limitations & trade-offs
- There is no early user–item interaction, so it is weak at fine-grained preferences. Mitigations: late-interaction models, multi-vector users, or generative retrieval such as TIGER.
- In-batch negatives are biased toward popular items. Without log-Q correction the model under-recommends head items.
- ANN indexes are stale until the next refresh, so fresh items need special handling.