RankMixerRankMixer: Scaling Up Ranking Models in Industrial Recommenders
A hardware-friendly Transformer variant: parameter-free token mixing instead of self-attention, and a separate FFN for every feature token.
RankMixer: Scaling Up Ranking Models in Industrial Recommenders§1The key contribution
Design the ranking backbone for GPUs: parameter-free token mixing plus per-token FFNs, scaling to billion-parameter models within the same latency.
Industrial ranking models are a patchwork of hand-designed interaction modules with low hardware utilization (MFU), so scaling parameters blows the latency budget. Self-attention is a poor fit for heterogeneous feature tokens.
Tokenize features into T semantically grouped tokens. Mix information across tokens with a parameter-free multi-head split-and-regroup, and give each token its own FFN so different feature subspaces don't compete for one shared FFN.
Figure 1. RankMixer architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
RankMixer turns hundreds of heterogeneous features into T tokens by grouping semantically related features and projecting each group to dimension D. Each block has two parts. Multi-head token mixing splits every token into H heads and regroups head h of all tokens into a new token h. This exchanges information across tokens with no parameters and no quadratic attention. Per-token FFNs then give each token its own FFN weights, because feature subspaces such as user, item and context are very different, and sharing one FFN lets the dominant ones crowd out the rest. The design scales to 1B+ parameters with high MFU, and a Sparse-MoE variant scales further.
Lineage. It borrows the MLP-Mixer idea (token mixing + channel MLP) and adapts it to recommendation features. It was deployed on Douyin's feed ranking. OneTrans (same org) later adds sequences into the same backbone with mixed parameterization.
§4Key equations
§5Why it works
- Large dense matmuls per token keep MFU high, which is essential when the latency budget is fixed and you want 100× more parameters.
- Self-attention assumes tokens live in a shared space. Feature tokens do not, which is why parameter-free mixing plus per-token weights works better here.
§6Limitations & trade-offs
- It needs careful manual feature grouping to form the tokens.
- It does not natively model long raw behavior sequences; those arrive pre-summarized. OneTrans addresses this.