OneTransOne Transformer for Feature Interaction and Sequence Modeling
Put behavior sequences and non-sequential features into one token sequence and run a single causal Transformer with mixed parameterization.
OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender§1The key contribution
Replace the separate 'sequence module' and 'feature-interaction module' with one causal Transformer over a unified token sequence.
Ranking models encode behavior sequences and other features in separate modules that meet only at the end. That blocks two-way information flow and makes it hard to scale either part with LLM-style techniques.
Tokenize sequential events (S-tokens) and other features (NS-tokens) into one sequence, S first and NS last, and stack causal Transformer blocks. S-tokens share weights while each NS-token gets its own. A pyramid schedule and cross-request KV caching keep it affordable.
Figure 1. OneTrans architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
Industrial rankers usually have two separate modules: a sequence encoder (DIN/SIM-style) that compresses user behaviors, and a feature-interaction network (DCN/RankMixer-style) over the rest. Information only meets at the end. OneTrans tokenizes both kinds of features into one sequence, with sequential S-tokens first and non-sequential NS-tokens last, and processes it with a stack of causal Transformer blocks. S-tokens share one set of QKV/FFN weights because they are homogeneous events. Each NS-token gets its own weights because each is a different feature subspace. A pyramid schedule keeps queries only for the most recent tokens at deeper layers, so the sequence shrinks until mostly the NS-tokens remain. Because S-tokens come first and attention is causal, their keys and values do not depend on the candidate. They can be computed once per request and reused across all candidates, and incrementally across requests (KV caching).
Lineage. It merges target-attention sequence modeling (DIN → SIM) and scalable feature interaction (DCN → Wukong → RankMixer) into one LLM-style backbone. Reported +5.68% per-user GMV in online A/B tests.
§4Key equations
§5Attention pattern
§6Why it works
- Unification lets behavior history and static features interact at every layer, not only at the final concat.
- Mixed parameterization balances homogeneous events (share weights, generalize across positions) against heterogeneous features (separate weights, as in RankMixer).
- It reuses LLM systems tricks: FlashAttention, mixed precision, activation recomputation and KV caching. That makes it the first such unified backbone practical in production.
§7Limitations & trade-offs
- Pyramid truncation and the cache only help if S-tokens precede NS-tokens. Candidate-specific features cannot influence how the sequence itself is encoded.
- Many token-specific NS parameters make it memory-heavy as the number of NS-tokens grows.