MINDMulti-Interest Network with Dynamic Routing
Represent each user with K interest vectors, extracted from behaviors by capsule-style dynamic routing, and retrieve with each one.
Multi-Interest Network with Dynamic Routing for Recommendation at Tmall§1The key contribution
Replace the single user vector of two-tower retrieval with several interest vectors, so a user who likes both hiking gear and baby products can retrieve both.
A two-tower user vector must sit near every item the user likes. With diverse interests, it ends up in a blurry average that matches none of them well, which hurts retrieval recall.
A multi-interest extractor uses behavior-to-interest (B2I) dynamic routing from capsule networks to cluster the user's behaviors into K interest capsules. Each capsule, combined with profile features, becomes a user vector. Training uses label-aware attention to pick the interest that matches the target item; serving runs one ANN query per interest.
Figure 1. MIND architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
MIND embeds a user's behaviors and applies a shared bilinear map. B2I dynamic routing then runs a few iterations: routing logits b_ij between behavior i and capsule j are softmaxed into coupling weights, each capsule is the squashed weighted sum of mapped behaviors, and the logits grow with agreement uⱼᵀ S eᵢ. The number of capsules K adapts to the history length. Each capsule is concatenated with profile features and passed through ReLU layers, giving K user vectors. In training, label-aware attention, softmax(pow(Vᵀ eᵢ, p)), weights the interests by their similarity to the target item, and a sampled softmax loss is applied. At serving, each of the K vectors queries the item ANN index. Deployed on Tmall's homepage.
Lineage. It extends YouTube-DNN/two-tower retrieval and borrows capsule routing from Sabour et al. (2017). ComiRec and other multi-interest retrievers followed, and the idea reappears in multi-vector user encoders for generative retrieval.
§4Key equations
§5Why it works
- Multiple vectors fix the averaging problem of a single embedding at almost no index cost, since items are still single vectors.
- Label-aware attention avoids the K interests collapsing into copies of one another.
§6Limitations & trade-offs
- Serving cost grows with K (several ANN queries per request), and the results must be merged.
- Routing iterations add latency, and the choice of K needs tuning.