DCN-v2Deep & Cross Network V2
Explicit, bounded-degree feature crosses learned by a stack of cheap cross layers.
DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems§1The key contribution
Learn bounded-degree feature crosses explicitly with a cheap cross layer, x₀ ⊙ (Wxₗ + b) + xₗ, instead of hoping an MLP finds them.
MLPs approximate multiplicative feature interactions surprisingly poorly, and hand-crafted crosses don't scale. DCN-v1's rank-1 cross layer was too restrictive for web-scale data.
Each cross layer multiplies the original input x₀ element-wise with a full-matrix linear map of the current state, which raises the polynomial degree by exactly one per layer. Low-rank and mixture-of-experts factorizations keep it affordable.
Figure 1. DCN-v2 architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
DCN learns feature crosses explicitly: each cross layer multiplies the original input x₀ element-wise with a linear transform of the current layer, then adds a residual. Stacking L layers yields polynomial interactions up to degree L+1. A plain MLP runs alongside (parallel) or on top (stacked) to capture implicit interactions.
Lineage. Successor of Wide & Deep (hand-crafted crosses) and DCN-v1 (rank-1 cross layer with a weight vector). v2 replaces the vector with a full matrix W and adds a low-rank mixture-of-experts variant (DCN-Mix) for efficiency.
§4Key equations
K low-rank experts with gates Gᵢ; cost drops from d² to roughly 2dr per expert.
§5Why it works
- Interaction degree grows exactly by one per layer, so depth directly controls the maximum cross order. That is easy to reason about.
- Cost per cross layer is one d×d matvec. That is far cheaper than enumerating pairwise interactions (FM-style) when d is large.
- The paper's empirical finding: MLPs alone are surprisingly bad at learning even 2nd-order dot products. Explicit cross layers fix that.
§6Limitations & trade-offs
- W is d×d, so it gets expensive when x₀ is very wide. That is why the low-rank DCN-Mix exists.
- It works on a single flattened vector, not on a set of feature tokens, so it does not scale as gracefully as the token-based backbones (Wukong, RankMixer).