xDeepFMeXtreme Deep Factorization Machine
A Compressed Interaction Network learns explicit high-order crosses at the vector level, alongside a linear part and a DNN.
xDeepFM: Combining Explicit and Implicit Feature Interactions for Recommender Systems§1The key contribution
Generate explicit, bounded-degree feature interactions at the vector-wise level (whole embeddings interact), not the bit-wise level used by DCN-v1 and MLPs.
DCN-v1's cross network interacts individual embedding dimensions (bit-wise) and is restricted to a special form, and MLPs learn interactions only implicitly. FM-style models explicitly interact whole field embeddings but only to 2nd order.
The Compressed Interaction Network (CIN) builds the k-th layer's feature maps from the Hadamard products of the previous layer's maps with the original field embeddings X⁰, compressed by learned filters like a CNN. Sum-pooling each layer gives explicit interactions of every order up to the depth. Combine it with a linear part and a DNN.
Figure 1. xDeepFM architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
xDeepFM embeds m fields into X⁰ ∈ ℝ^{m×D}. In the CIN, layer k+1 takes the Hadamard product of every row of X^k (H_k maps) with every row of X⁰ (m fields), giving an H_k × m × D tensor. Learned weights W^{k,h} compress it into H_{k+1} new maps, similar to a convolution along the embedding dimension. Each CIN layer's maps are sum-pooled over D and concatenated. A plain linear model and a DNN run in parallel, and the three outputs are summed and passed through a sigmoid. CIN gives explicit vector-wise interactions up to order T+1 with interpretable structure.
Lineage. It sits between DeepFM (2nd-order FM) and DCN (bit-wise crosses). DCN-v2's full-matrix cross layer later closed much of the expressiveness gap at lower cost.
§4Key equations
§5Why it works
- Vector-wise crosses respect field boundaries, which tends to generalize better than bit-wise ones on sparse data.
- The explicit order structure (order = depth + 1) makes the model easier to analyze than a pure MLP.
§6Limitations & trade-offs
- CIN is expensive, O(m·H²·D·T), and was often too slow for production compared with DCN-v2.
- The extra complexity often brought only small gains over well-tuned DCN or DeepFM baselines.