OneRecEnd-to-End Generative Recommendation
One encoder–decoder with MoE replaces the whole retrieve → pre-rank → rank cascade, generating semantic IDs of videos and aligned with RL.
OneRec Technical Report§1The key contribution
Collapse the multi-stage recommendation cascade into a single generative model that directly decodes which videos to show, and make it cheaper to run than the cascade it replaces.
Industrial recommenders are cascades of separately trained models (retrieval, pre-ranking, ranking) with inconsistent objectives. Most of the compute goes to communication and storage rather than model FLOPs, and the ranking models run at very low GPU utilization (MFU), so they cannot benefit from scaling.
Tokenize videos into collaborative-aware multimodal semantic IDs. Encode the user with four behavior pathways, from static profile to 100k-item lifelong history. Decode the next videos' IDs with a sparse-MoE decoder. Pretrain with next-token prediction, then optimize for user value with a learned reward and an RL algorithm (ECPO).
Figure 1. OneRec architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
OneRec is Kuaishou's end-to-end generative recommender. Each video is tokenized into a 3-level semantic ID: a multimodal LLM summarizes content, a QFormer compresses it, contrastive alignment with collaboratively similar pairs adds behavior signal, and RQ-KMeans quantizes the result. The encoder concatenates four user pathways (static, short-term, positive feedback, lifelong) with positional embeddings and applies Transformer layers. The decoder takes [BOS] plus the target video's codewords, applies causal self-attention, cross-attention to the encoder, and sparse MoE feed-forward layers, and is trained with next-token prediction on billions of samples per day. Post-training combines rejection-sampling fine-tuning with RL: a reward system (P-Score preference model, format reward, business reward) and ECPO, an early-clipped variant of GRPO. At inference, beam search decodes semantic IDs directly into the videos to show, with no separate retrieval and ranking stages.
Lineage. It extends TIGER-style semantic-ID generation from retrieval to the whole pipeline. The original OneRec (Feb 2025) introduced session-wise generation and iterative DPO-based preference alignment. The technical report (Jun 2025) scaled it up and replaced DPO with ECPO.
§4Key equations
§5Why it works
- End-to-end generation removes the mismatch between cascade stages: the same model decides recall and order.
- Moving compute from feature lookups and communication into dense Transformer FLOPs is what made scaling pay off (MFU 5× higher).
- RL with a learned preference reward lets the model optimize what the business wants (stay time, lifetime), not only imitate logged behavior.
§6Limitations & trade-offs
- Huge infrastructure: dozens of 8-GPU servers for training, a separate inference service for RL sampling, custom TensorRT kernels.
- Semantic-ID collisions and codebook re-training complicate catalog churn.
- It only partly replaced the cascade (~25% of QPS at report time); the reward model can be gamed if not carefully constrained.