DLRMDeep Learning Recommendation Model
Dense features through a bottom MLP, sparse features through embedding tables, all pairwise dot products, then a top MLP.
Deep Learning Recommendation Model for Personalization and Recommendation Systems§1The key contribution
A deliberately simple, hardware-aware reference design: embeddings + bottom MLP, pairwise dot products, top MLP.
By 2019 recommendation models dominated datacenter compute at Meta, but there was no standard, well-engineered architecture to benchmark systems or hardware against.
Combine factorization-machine-style dot-product interactions with MLPs in the simplest effective form, and co-design the parallelism: model-parallel embedding tables, data-parallel MLPs.
Figure 1. DLRM architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
DLRM is the reference ranking architecture behind Meta's ads and feeds. Dense features are projected by a bottom MLP to the same dimension d as the embeddings. Each sparse feature looks up a row in its own embedding table. The interaction layer takes the dot product of every pair of these vectors (like a factorization machine), concatenates the results with the dense vector, and a top MLP produces the CTR. It was designed for hardware: embedding tables are sharded across devices (model parallel) while the MLPs are data parallel.
Lineage. Factorization machines plus MLPs, simplified for scale. Many 'DLRM-style' models swap the interaction layer for DCN, Transformers or Wukong while keeping the same sparse/dense skeleton.
§4Key equations
§5Why it works
- It is deliberately simple. Interactions are only dot products, which are very cheap. Almost all the parameters live in embedding tables.
- Its systems design (sharded tables + all-to-all, data-parallel MLPs) shaped how recommendation training is still done: TorchRec, FBGEMM.
§6Limitations & trade-offs
- Interactions are only 2nd order, and the compute per example is tiny. That makes it memory-bound and hard to scale with more FLOPs. Wukong and HSTU address that.
- All features must share one dimension d.