MTGRMeituan Generative Recommendation
An HSTU-style Transformer ranker that keeps the old DLRM cross features, and scores all of a user's candidates in a single sequence.
MTGR: Industrial-Scale Generative Recommendation Framework in Meituan§1The key contribution
Get generative-recommender scaling without discarding the hand-crafted cross features that DLRMs depend on: put the candidates and their cross features into the token sequence.
Generative recommenders such as HSTU drop most engineered features, especially user-item cross features. In Meituan's delivery setting, losing them hurt quality more than scaling could recover. Per-candidate Transformer inference is also far too expensive.
Build one sequence per user: profile tokens, long history, real-time history, then K candidate tokens, each carrying its item features plus cross features. A dynamic attention mask prevents leakage, Group LayerNorm handles each semantic group separately, and one pass ranks all K candidates.
Figure 1. MTGR architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
MTGR ranks restaurant and food candidates at Meituan. Each training sample is one user: profile features U (one token each), a long-term history S (up to ~1000 items), a real-time sequence R (recent hours, up to ~100), and K candidate tokens. Each candidate token concatenates the item's features with the user-item cross features that the production DLRM used. Every feature or item is projected to d_model. Stacked HSTU-style blocks apply Group LayerNorm, project to Q, K, V, U, compute SiLU attention under a dynamic mask, gate with U, and project back with a residual. The mask makes static tokens globally visible and real-time tokens causal, and lets each candidate see the user context but no other candidate. CTR/CTCVR heads read each candidate's final token, so the whole request is scored in one pass.
Lineage. It builds directly on HSTU (Meta) but reverses its 'drop features' stance. Its sequence design is a close cousin of OneTrans: both put candidate-specific tokens at the end of a user-level sequence.
§4Key equations
§5Attention pattern
§6Why it works
- Cross features encode years of domain knowledge in a form a sequence model cannot easily rediscover; keeping them made scaling additive rather than a replacement.
- Batching candidates as tokens with a self-only mask gives Transformer-quality ranking at roughly DLRM serving cost.
§7Limitations & trade-offs
- Engineered cross features still have to be computed and maintained, so it is less 'end-to-end' than HSTU or OneRec.
- The sequence grows with K, and K is bounded by the latency budget.