SASRecSelf-Attentive Sequential Recommendation
A decoder-only (GPT-style) Transformer over the user's item sequence, trained to predict the next item at every position.
Self-Attentive Sequential Recommendation§1The key contribution
Model the user's sequence with a causal (GPT-style) self-attention Transformer and train next-item prediction at every position.
Markov-chain models only look one step back, and RNN models (GRU4Rec) are slow and struggle with long-range dependencies, especially on dense datasets.
Apply a decoder-only Transformer to item IDs. Attention picks the relevant past items adaptively, and causal masking turns each sequence into t training examples in one parallel pass.
Figure 1. SASRec architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
SASRec treats a user's history i₁, i₂, …, i_t as a sentence. Item embeddings plus learned position embeddings pass through b causal self-attention blocks, so position k can only attend to items 1…k. The output at position k is matched against all item embeddings by dot product to predict i_{k+1}. Every position is a training example, which makes training efficient. At inference the last hidden state F_t is the user representation.
Lineage. It replaced RNN (GRU4Rec) and CNN (Caser) sequence models. BERT4Rec is the bidirectional, masked-item variant. HSTU and generative recommenders scale the same idea with actions and much longer contexts.
§4Key equations
§5Attention pattern
§6Why it works
- Attention adapts how far back to look: on sparse datasets it focuses on the last few items, on dense ones it uses long-range context.
- Weight tying between input and output item embeddings reduces parameters and helps generalization.
§7Limitations & trade-offs
- Only item IDs and positions. No side features, timestamps or action types (HSTU adds those).
- Quadratic in sequence length, so practical histories are limited to a few hundred items without tricks.