DIENDeep Interest Evolution Network
Extract a latent interest at every step with a GRU, then let the interests relevant to the candidate evolve through an attention-gated GRU.
Deep Interest Evolution Network for Click-Through Rate Prediction§1The key contribution
Model interests as a hidden process that evolves over time, rather than as a bag of past behaviors, and let the candidate steer which evolution path matters.
DIN treats behaviors as independent items and ignores their order. Raw behaviors are also only a noisy proxy for the user's underlying interest, which drifts over time and differs across interest types.
Two layers. An interest extractor GRU turns behaviors into interest states, made more discriminative by an auxiliary loss that asks each state to predict the next behavior. An interest evolving layer then runs an AUGRU, a GRU whose update gate is scaled by attention to the candidate, so only relevant interests drive the final state.
Figure 1. DIEN architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
DIEN embeds the behavior sequence and runs a GRU (the interest extractor) to get hidden states h₁…h_T. An auxiliary loss trains each hₜ to distinguish the real next behavior e_{t+1} from a sampled negative, which gives every step its own supervision. Next, attention scores aₜ between each hₜ and the target ad are computed. A second GRU, the interest evolving layer, uses AUGRU (GRU with Attentional Update gate): its update gate is scaled by aₜ, so steps unrelated to the target barely change the evolving state. The final state h′_T is concatenated with the target ad, user profile and context, and an MLP predicts CTR. Deployed in Taobao display advertising, it gave +20.7% CTR over the then-serving model.
Lineage. It directly extends DIN (target attention) with sequence modeling. Its successors DSIN (session-based) and SIM (lifelong sequences) push further on structure and length.
§4Key equations
§5Why it works
- Interest drift is real in e-commerce: a user browsing coats in October is not the same user in May. Sequential modeling captures that.
- Gating the update (AUGRU) rather than the input (AIGRU) or the output (AGRU) works best, because it controls how much each step can change the state.
§6Limitations & trade-offs
- GRUs are sequential, so they are slow to train and serve on long histories (practical T ≈ 50).
- The evolving GRU depends on the candidate, so it must be recomputed per candidate, which is expensive.