Wide & DeepWide & Deep Learning
A linear model over hand-crafted crosses for memorization, plus an MLP over embeddings for generalization, trained jointly.
Wide & Deep Learning for Recommender Systems§1The key contribution
Train a memorizing linear model and a generalizing deep network jointly, so each covers the other's failure mode.
Linear models over cross-product features memorize well but cannot generalize to unseen feature pairs. Embedding-based deep models generalize but over-generalize for niche users, recommending loosely related items.
Sum the logits of a wide (linear + crosses) model and a deep (embeddings + MLP) model and train both with one loss. The wide part only needs to fix the deep part's mistakes, so it can stay small.
Figure 1. Wide & Deep architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
The wide part is a generalized linear model over raw sparse features and engineered cross-product features such as AND(installed_app=X, impression_app=Y). It memorizes specific co-occurrences that worked before. The deep part embeds the sparse features, concatenates them with the dense features and runs an MLP, which generalizes to unseen combinations. Both logits are summed and trained jointly. It was the first widely deployed deep ranking model (Google Play).
Lineage. It bridged classic logistic regression with feature crosses and neural nets. DeepFM, DCN and DLRM all exist to remove the hand-engineered wide part by learning the crosses automatically.
§4Key equations
§5Why it works
- Memorization and generalization fail in different ways. Combining both gives relevance (wide) and diversity (deep).
- Joint training lets the wide part stay small: it only has to fix what the deep part gets wrong.
§6Limitations & trade-offs
- The crosses are hand-engineered. Every new interaction needs a human, and the feature space grows combinatorially.
- The deep part learns multiplicative interactions inefficiently. Later models (DeepFM, DCN) make the crosses explicit and learned.