DINDeep Interest Network
Pool the user's behavior history with attention weights computed against the candidate ad, so the user vector depends on what is being scored.
Deep Interest Network for Click-Through Rate Prediction§1The key contribution
Let the candidate decide which parts of the user's history matter: attention-weighted pooling of behaviors conditioned on the target ad.
Embedding-and-MLP models pool a user's whole history into one fixed vector, the same for every candidate. A user with diverse interests gets a blurry representation that limits CTR accuracy.
Compute a relevance weight between each past behavior and the candidate with a small 'activation unit', then take a weighted (not softmax-normalized) sum. The user representation becomes candidate-specific: local activation.
Figure 1. DIN architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
A user who bought a coat, sneakers and a phone case has different relevant interests depending on the ad. Earlier models averaged all behavior embeddings into one fixed user vector. DIN runs a small 'activation unit' MLP on each (behavior, candidate) pair to get a relevance weight wᵢ, then sums the behaviors weighted by wᵢ. The weights are not softmax-normalized, so the magnitude of the pooled vector still reflects how strong the interest is.
Lineage. This is target attention, the template for most behavior-sequence models in ranking. DIEN adds a GRU to model interest evolution. SIM and TWIN extend it to 10k+-long histories with a retrieval step. OneTrans folds it into a single transformer.
§4Key equations
§5Why it works
- A fixed-length user vector has to compress every interest. Conditioning on the candidate means it only needs to surface the relevant one.
- Dropping softmax keeps 'how much' as well as 'which', which matters for CTR calibration.
- Cost is linear in history length T per candidate, fine for T ≈ 50–1000.
§6Limitations & trade-offs
- Each behavior is scored independently: no ordering and no interaction between behaviors. DIEN and transformers fix that.
- The cost is per candidate, and it grows with T × #candidates. Very long sequences need a retrieval step (SIM) or caching (OneTrans).