WukongWukong: Towards a Scaling Law for Large-Scale Recommendation
Stack factorization-machine blocks so that interaction order grows exponentially with depth, and scale it like an LLM.
Wukong: Towards a Scaling Law for Large-Scale Recommendation§1The key contribution
Stack factorization-machine blocks so interaction order grows exponentially with depth, giving recommendation its first clean scaling law.
Upscaling DLRM-style models (wider MLPs, bigger embeddings) yields diminishing returns. Their interaction modules do not convert extra compute into quality.
Treat embeddings as an n×d matrix and repeat one layer: an FM block (pairwise products → MLP → new embeddings) in parallel with a linear compress block, plus residual and LayerNorm. Depth gives 2ˡ-order interactions; width and depth scale smoothly.
Figure 1. Wukong architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
DLRM-style models stop improving when you add compute, because their interaction layer is shallow. Wukong keeps the embeddings-as-tokens view (X ∈ ℝ^{n×d}) and stacks identical layers. Each layer has a Factorization Machine Block (FMB), which computes pairwise interactions XXᵀ and turns them back into embeddings with an MLP, and a Linear Compress Block (LCB), which linearly recombines the input embeddings. Their outputs are concatenated, added to the residual and layer-normalized. Because each layer's output feeds the next FM, layer i captures interactions up to order 2ⁱ. The paper shows a scaling law over two orders of magnitude of compute.
Lineage. A direct successor to DLRM and DCN in spirit, aimed at scaling. RankMixer and OneTrans are later 'scalable backbones' that take a more Transformer-like route.
§4Key equations
§5Why it works
- Binary-exponentiation-style stacking reaches high-order interactions with few layers.
- It is the first recsys architecture with a clean scaling law (quality vs. GFLOP/example) over about 100× compute, the regime where LLM-style investment pays off.
§6Limitations & trade-offs
- It focuses on feature interaction only. Long behavior sequences still need a separate sequence module, which OneTrans unifies.
- Pairwise FM is O(n²d) per layer, so it needs the low-rank optimization at large n.