GenPageEnd-to-End Generative Homepage Construction
One decoder-only Transformer reads the user's context as a prompt and writes the whole multi-row homepage, rows and titles, token by token.
GenPage: Towards End-to-End Generative Homepage Construction at Netflix§1The key contribution
Replace the multi-stage page-construction stack with a single model that generates the entire structured homepage autoregressively, trained like an LLM: pretraining, then post-training.
A streaming homepage is normally built by a cascade: candidate generation, per-row rankers, row selection and ordering, then business-rule layers and dedup. Each stage optimizes its own proxy, nobody optimizes the page as a whole, and the stack is slow and hard to change.
Treat 'user + request context' as a prompt and 'the page' as the response. A domain-specific tokenizer gives every title and every row type its own token. The model generates row token → entities → next row token … left-to-right, top-to-bottom. Constrained decoding enforces business rules. Pretrain on engaged production pages, then post-train with weighted binary classification or RL against a page-level reward.
Figure 1. GenPage architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
GenPage frames homepage construction as conditional sequence generation. The prompt holds the member's profile, request context (device, time) and recent history as discrete tokens. The response is the homepage written top-to-bottom: a row token such as 'Korean TV Shows', then the entity tokens in that row, then the next row token, and so on. A ~200M-parameter decoder-only Transformer is pretrained with next-token prediction on production pages that received positive engagement. It is then post-trained either with Weighted Binary Classification, where each token's logit becomes a value estimate trained on reward-derived labels, or with RL (Dr. GRPO) against a page-level reward model. Constrained decoding applies business rules, and hybrid row decoding fills the tail of each row in a single pass. Online, the WBC variant beat a mature multi-stage production system (+0.24% core engagement, p<0.001) while cutting end-to-end latency by 20%.
Lineage. It applies the LLM recipe (tokenize → pretrain → post-train with preference/RL) to whole-page recommendation. It is closest to OneRec's 'session-wise' generation, but its output is a structured 2-D page with row semantics. Its sibling at Netflix is GenRec, an LLM-backed ranker.
§4Key equations
Sketch: zᵧ is the token's logit, rₜ the reward attributed to it; random negative tokens are added as zero-label samples. See the paper for the exact formulation.
§5Why it works
- The model sees the entire page it has generated so far, so it can balance rows against each other (diversity, no duplicate titles) instead of each row being ranked in isolation.
- It shows that information is the bottleneck: enriching the prompt beat 7× more parameters, a useful prior for anyone scaling generative recommenders.
- Replacing the cascade with one model reduced latency, the opposite of the usual worry about generative serving cost.
§6Limitations & trade-offs
- Long histories are still summarized by hand before tokenization, which partly reintroduces feature engineering ('prompt engineering').
- Online gains of the shipped variant are modest (+0.24%) next to a heavily tuned production stack; the RL variant's gains are mainly reported offline.
- Each new entity needs a vocabulary slot and embedding, so catalog churn requires vocabulary management.