BERT4RecSequential Recommendation with Bidirectional Encoder Representations from Transformer
BERT for item sequences: mask random items and predict them using context from both directions.
BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer§1The key contribution
Model user sequences bidirectionally by training with a Cloze (masked-item) objective instead of left-to-right next-item prediction.
Left-to-right models (SASRec, GRU4Rec) only let each item see its past. Real behavior sequences are not strictly ordered: the order of nearby clicks is often arbitrary, so context from both sides helps.
Use a bidirectional Transformer encoder. During training, randomly replace items with [mask] and predict them from both sides. At inference, append a [mask] at the end of the sequence and predict what fills it.
Figure 1. BERT4Rec architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
BERT4Rec embeds each item plus a learned position embedding and stacks L Transformer encoder layers with full (bidirectional) multi-head self-attention, position-wise GELU FFNs, residuals and LayerNorm. Training uses the Cloze task: a random subset of items is replaced by [mask], and the model predicts the original item at each masked position with a softmax over all items, using the transposed embedding matrix as the output layer. Some sequences are trained with only the last item masked to match inference. At serving time, [mask] is appended after the user's last interaction and its prediction gives the next-item ranking.
Lineage. It applies BERT (2018) to recommendation as the bidirectional counterpart of SASRec. Later re-evaluations found SASRec with a full softmax loss is often just as strong. The difference was largely the loss, not the bidirectionality.
§4Key equations
§5Attention pattern
§6Why it works
- Masking yields many targets per sequence (like BERT), which acts as data augmentation and regularization.
- Bidirectional context helps when local order is noisy, e.g. browsing several items in one session.
§7Limitations & trade-offs
- The training/inference mismatch (random masks vs one trailing mask) needs mitigation.
- Full-softmax training over a large catalog is expensive. Later studies show much of its edge over SASRec came from the loss, not the architecture.