SIMSearch-based Interest Model
Two-stage attention over lifelong histories: a cheap search keeps only the behaviors related to the candidate, then precise attention runs on those.
Search-based User Interest Modeling with Lifelong Sequential Behavior Data for Click-Through Rate Prediction§1The key contribution
Make target attention scale to lifelong histories (tens of thousands of behaviors) by first searching the history for candidate-relevant behaviors, then attending only to those.
Longer behavior histories keep improving CTR, but DIN/DIEN-style attention costs O(T) per candidate. At T ≈ 50,000 and hundreds of candidates, that is impossible under a ~10 ms latency budget, and memory-network approaches such as MIMN blur interests in a fixed-size memory.
Cascade the search. A General Search Unit (GSU) cheaply selects the top-K behaviors related to the candidate, either by exact category match (hard search) or by embedding inner product (soft search). An Exact Search Unit (ESU) then applies time-aware multi-head target attention to just those K.
Figure 1. SIM architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
SIM handles user histories of up to ~54,000 behaviors in Alibaba display advertising. Stage 1, the General Search Unit, filters the lifelong sequence down to the K behaviors relevant to the candidate. The hard version keeps behaviors with the same category, served from an offline user-behavior tree. The soft version uses inner products between projected behavior and candidate embeddings. Stage 2, the Exact Search Unit, concatenates each selected behavior's embedding with a time-gap embedding and runs multi-head target attention with the candidate as the query. The result joins the usual short-term features in the CTR MLP. Deployed in 2019, it gave +7.1% CTR and +4.4% RPM.
Lineage. It extends DIN to lifelong scale and spawned the 'retrieve-then-attend' family (UBR4CTR, ETA, SDIM, TWIN). OneTrans and HSTU take a different route: efficient attention over longer raw sequences.
§4Key equations
§5Why it works
- Most of a lifelong history is irrelevant to any single candidate, so a cheap filter loses little and saves orders of magnitude of compute.
- Plain category matching was nearly as good as learned soft search and far easier to serve, which is why it shipped.
§6Limitations & trade-offs
- The GSU and ESU use different relevance measures, so stage 1 can drop behaviors stage 2 would have liked. TWIN later aligns the two.
- Hard search depends on a good category taxonomy.