TIGERTransformer Index for GEnerative Recommenders
Give every item a short tuple of semantic codewords, then have a seq2seq Transformer generate the next item's code directly.
Recommender Systems with Generative Retrieval§1The key contribution
Turn retrieval into generation: give every item a semantic ID (a short code tuple) and have a seq2seq Transformer generate the next item's code.
Dual-encoder retrieval needs an embedding per item, a separate ANN index, and handles cold-start poorly. Atomic item IDs give the model no notion of item similarity.
Quantize item content embeddings with an RQ-VAE into hierarchical codewords, so similar items share prefixes. Then train an encoder–decoder to autoregressively decode the next item's codewords from the history. The index lives in the model's weights.
Figure 1. TIGER architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
Instead of scoring every item by dot product, TIGER makes the model generate the next item. First, each item's content embedding is quantized with an RQ-VAE into a semantic ID: a tuple such as (5, 23, 55), where items with similar content share prefixes. A user's history becomes a token sequence of these codes. An encoder–decoder Transformer reads the history and decodes the next item's code token by token with beam search. The 'index' lives in the model's weights, and new items with known content get meaningful IDs immediately.
Lineage. Inspired by DSI (Differentiable Search Index) for document retrieval. It opened the line of 'semantic ID' generative recommenders. Many 2024–25 industrial systems (e.g. OneRec) use the same tokenization idea.
§4Key equations
§5Why it works
- The item vocabulary shrinks from millions of IDs to D × K codewords, so the embedding table no longer grows with the catalog.
- Similar items share prefixes, so knowledge transfers to cold-start items with no interactions.
- Beam search over a hierarchy works like learned tree-based retrieval.
§6Limitations & trade-offs
- Inference is autoregressive decoding with beam search, which costs more than a single ANN lookup.
- Semantic IDs come from content only unless collaborative signals are injected, and re-training the RQ-VAE reshuffles the IDs.