HLLMHierarchical Large Language Model
Two stacked LLMs: one reads an item's text and outputs its embedding, the other reads a sequence of those embeddings and predicts the next one.
HLLM: Enhancing Sequential Recommendations via Hierarchical Large Language Models for Item and User Modeling§1The key contribution
Split the recommender into an Item LLM and a User LLM, both initialized from pretrained LLMs, so world knowledge helps item understanding while sequences stay short (one embedding per item).
Feeding whole user histories as text into one LLM makes sequences enormous and slow. Pure ID-based sequential models cannot use the knowledge in item text and handle cold items poorly.
Item LLM: item text + a special [ITEM] token → the final hidden state is the item embedding. User LLM: a causal LLM over item embeddings (its text vocabulary is removed) predicts the next item's embedding. Train both end-to-end with a generative next-embedding loss, then optionally fine-tune a discriminative ranking head.
Figure 1. HLLM architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
HLLM uses two LLMs in a hierarchy. The Item LLM takes an item's text (title, tags, description) with a special [ITEM] token appended; that token's final hidden state is the item embedding. The User LLM is another pretrained decoder whose word embeddings are dropped. Its inputs are the sequence of item embeddings from the user's history, and at each position it outputs a predicted embedding for the next item. Training is end-to-end: a generative loss (InfoNCE between predicted and true next-item embeddings, with sampled negatives) for retrieval, and a discriminative loss (target item fused with the user sequence, BCE) for ranking. Initializing both from pretrained LLMs and fine-tuning everything is what gives the gains, and quality keeps improving at 7B scale.
Lineage. It sits between ID-based sequential models (SASRec, HSTU) and text-only LLM recommenders (P5, TALLRec). Later systems often keep its idea of an LLM item encoder feeding a sequence model.
§4Key equations
§5Why it works
- The hierarchy keeps sequence length equal to the number of items, which is the main reason LLM-scale models become trainable on long histories.
- Item embeddings come from text, so new items get good representations immediately.
- Pretrained weights transfer: language-model knowledge helps recommendation once both levels are fine-tuned.
§6Limitations & trade-offs
- Two LLMs is a lot of compute; item embeddings must be cached and refreshed for serving.
- It is text-only in the original paper. Visual content and fine-grained collaborative IDs need extra signals.