HSTUHierarchical Sequential Transduction Units
Reformulate retrieval and ranking as sequential transduction over (item, action) tokens, with an attention block redesigned for recommendation data.
Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations§1The key contribution
Recast recommendation as generative sequential transduction over raw user actions, with an attention block redesigned for recsys data, and show it follows scaling laws.
DLRMs rely on thousands of hand-engineered features and stop improving as compute grows. Standard Transformers handle recommendation's huge, non-stationary vocabularies and very long histories poorly.
Represent each user as one chronological stream of (item, action) tokens and train retrieval and ranking as next-token tasks. Replace attention + FFN with HSTU: pointwise SiLU attention with time-aware relative bias, plus a gated pointwise transform.
Figure 1. HSTU architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
Generative Recommenders (GRs) drop most hand-engineered features. A user's entire history becomes one chronological sequence of item tokens Φᵢ interleaved with action tokens aᵢ (click, like, skip…). Retrieval predicts the next item; ranking predicts the action at an item position. Each HSTU layer replaces the Transformer's attention + FFN with three steps: a pointwise projection producing U, V, Q, K; a spatial aggregation that uses pointwise SiLU attention instead of softmax, with relative position and time biases; and a gated pointwise transformation. Quality follows a power-law in training compute up to 1.5 trillion parameters, and the model is 5–15× faster than FlashAttention2 Transformers on 8k-length sequences.
Lineage. It extends SASRec-style causal sequence models to recommendation with heterogeneous features and actions. The scaling-law results pushed the field toward 'foundation-model' rankers (OneTrans, RankMixer, MTGR…).
§4Key equations
§5Attention pattern
§6Why it works
- Removing softmax keeps engagement intensity, and the authors found it more robust to non-stationary, ever-growing vocabularies.
- Collapsing attention + FFN into a single fused block reduces activation memory, which allows much deeper models within the same HBM.
- Stochastic Length (randomly shortening long histories in training) and M-FALCON (micro-batched candidate scoring with KV reuse) make it cheap enough to deploy.
§7Limitations & trade-offs
- It moves away from hand-engineered features. Teams with mature feature pipelines must re-think how to inject them.
- Training cost is large, and gains follow scaling laws, so small models do not beat strong DLRM baselines by much.