AutoIntAutomatic Feature Interaction Learning via Self-Attentive Neural Networks
Treat each feature field as a token and let multi-head self-attention decide which fields should interact.
AutoInt: Automatic Feature Interaction Learning via Self-Attentive Neural Networks§1The key contribution
Use multi-head self-attention over feature-field embeddings to learn which high-order interactions matter, explicitly and interpretably.
FM-style models enumerate interactions with fixed forms (all pairs, equal treatment), while MLPs learn them implicitly and opaquely. Neither tells you which combinations the model relies on.
Embed every field (categorical and numerical) into the same d-dimensional space, then stack 'interacting layers', which are multi-head self-attention plus a residual. Each layer combines fields weighted by learned relevance, so stacking L layers models interactions of increasing order, and the attention maps can be inspected.
Figure 1. AutoInt architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
AutoInt maps each of the M input fields to a d-dimensional embedding. Categorical fields use embedding lookup, and numerical fields multiply a learned vector by the scalar value. An interacting layer applies multi-head self-attention across the M field embeddings: for each head, field m's new representation is Σₖ α_{m,k} W_V eₖ, with α given by a softmax over ⟨W_Q e_m, W_K eₖ⟩. Heads are concatenated, a projected residual W_Res e_m is added, and a ReLU applied. L such layers are stacked, and the final field representations are concatenated and passed to a logistic output. It matched or beat DeepFM, xDeepFM and DCN with fewer parameters, and adding an MLP branch (AutoInt+) helps further.
Lineage. It is the first widely cited use of Transformer-style self-attention for CTR feature interaction. Its 'fields as tokens' view underlies later backbones such as RankMixer, Wukong and HSTU-style rankers.
§4Key equations
§5Why it works
- Attention gives a different weight to each pair of fields, which fixed-form FM interactions cannot do.
- The 'fields as tokens' framing let recommendation reuse Transformer machinery and, later, Transformer scaling.
§6Limitations & trade-offs
- Softmax attention assumes fields share a semantic space. RankMixer later argued this mismatch limits pure self-attention on heterogeneous features.
- Cost is O(M²) in the number of fields per layer.