MMoEMulti-gate Mixture-of-Experts
Shared experts, with one softmax gate per task deciding how to mix them.
Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts§1The key contribution
Share a pool of experts across tasks but give every task its own gate, so the model learns how much to share.
Shared-bottom multi-task models force every objective (click, watch time, like…) through one representation. When tasks are weakly related or conflict, quality drops (negative transfer).
Replace the shared bottom with n experts and add one softmax gate per task. Each task gets its own input-dependent mixture of the experts, so related tasks can share experts and conflicting ones can specialize.
Figure 1. MMoE architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
Ranking usually predicts several objectives at once: click, like, share, watch time. A shared-bottom network forces all tasks to use the same representation, which hurts when tasks conflict. MMoE replaces the shared bottom with n expert networks. Each task has its own gate, a softmax over experts computed from the input, and gets its own weighted mixture. Related tasks learn to share experts, and conflicting tasks learn to use different ones.
Lineage. It generalizes One-gate MoE (OMoE) and shared-bottom models. PLE adds task-specific experts to reduce the 'seesaw' effect. YouTube's 2019 ranking system uses MMoE.
§4Key equations
§5Why it works
- The gates make the amount of sharing learnable. Highly correlated tasks converge to similar gates, and weakly correlated ones diverge.
- Gating adds few parameters (n × d per task), so it is much cheaper than separate models per task.
§6Limitations & trade-offs
- Gates can collapse onto a few experts (polarization), and all experts are still shared, so negative transfer remains. PLE addresses this.
- Balancing the loss weights across tasks is still a manual or heuristic problem.