GenRec (JD)A Preference-Oriented Generative Framework
A decoder-only LLM over semantic IDs that is trained to generate a whole page of interactions, then aligned with RL.
GenRec: A Preference-Oriented Generative Framework for Large-Scale Recommendation§1The key contribution
Make semantic-ID generative retrieval practical at e-commerce scale with three fixes: supervise pages instead of single items, compress the prompt, and align with a stabilized RL objective.
Point-wise next-item training is ambiguous when one request page has several clicked or ordered items: there are many correct 'next items'. Multi-token semantic IDs also triple the prompt length, and naïve RL on preference rewards over-optimizes and produces invalid IDs.
Use a Qwen2.5 decoder over RQ-KMeans semantic IDs. Train it to emit the full page of items a user interacted with, ordered by intensity (order > cart > click). Merge each item's 3 SID tokens into one vector on the input side only. Post-train with GRPO plus an NLL anchor (GRPO-SR) and gated rewards.
Figure 1. GenRec (JD) architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
JD's GenRec is a decoder-only generative retriever deployed in the JD app. Items are tokenized into 3-level semantic IDs by running RQ-KMeans on multimodal embeddings from Qwen2.5-VL. The user's history is a sequence of SIDs. An asymmetric linear Token Merger compresses each item's three SID tokens into one input vector, which halves the prompt without touching the output side. A Qwen2.5 backbone (3B, deep and narrow) is fine-tuned with page-wise NTP: the target is the list of items the user ordered, carted or clicked on that page, sorted by intensity. It is then aligned with GRPO-SR, which combines GRPO on gated preference rewards with an NLL regularizer on real positives. Serving uses standard beam search. A month-long A/B test reported +9.5% clicks and +8.7% transactions over the existing retrieval pipeline, with larger gains on long-tail items.
Lineage. It builds on TIGER-style semantic IDs and decoder-only LLM recommenders, and uses DeepSeek-style GRPO for alignment. It is not the same model as Netflix's GenRec, which is a discriminative LLM ranker with an ID scoring head.
§4Key equations
§5Why it works
- Page-level targets fit e-commerce, where one request page leads to several interactions, better than single next-item targets.
- Asymmetric compression (merge inputs, keep outputs) is a cheap, general trick for any multi-token-ID generator.
- The NLL anchor works like a KL constraint grounded in real data, guarding RL against reward hacking.
§6Limitations & trade-offs
- It relies on an external preference model (SIM-based) for rewards, so alignment can be no better than that scorer.
- Beam search over 3-token IDs still costs more per request than a single ANN lookup.