PLEProgressive Layered Extraction
Separate task-specific and shared experts explicitly, and stack the extraction across layers.
Progressive Layered Extraction (PLE): A Novel Multi-Task Learning (MTL) Model for Personalized Recommendations§1The key contribution
Explicitly separate task-specific experts from shared experts, and refine the separation progressively over several layers.
Even with MMoE, improving one task often hurts another (the 'seesaw' phenomenon), because every expert is shared and gradients from different tasks interfere.
Give each task private experts that only its own gate can use, keep a shared expert group, and stack these Customized Gate Control (CGC) units so higher layers progressively extract task-specific semantics.
Figure 1. PLE architecture. Arrows show data flow; annotations show tensor shapes or symbols. Click a block for details, or use the walkthrough to step through the forward pass.
§2Breaking it down
The contribution, piece by piece. Select a card to highlight the blocks it refers to in Figure 1.
§3How it works
MMoE often shows a 'seesaw' effect: improving one task hurts another, because every expert is shared. PLE splits the experts into task-specific groups and a shared group. In a Customized Gate Control (CGC) unit, task A's gate mixes only A's experts and the shared experts, and likewise for B. A shared gate mixes everything to pass a shared representation to the next layer. Stacking several such extraction layers separates task-specific and shared knowledge progressively.
Lineage. It builds on MMoE. A single-layer PLE equals CGC. Widely used in Tencent Video and other industrial multi-task rankers.
§4Key equations
§5Why it works
- Explicit task-specific experts guarantee each task some capacity the other task's gradients cannot overwrite. This reduces the seesaw effect.
- Gate analysis in the paper shows expert usage diverging across layers, which is evidence of real separation.
§6Limitations & trade-offs
- More hyperparameters: expert counts per group, number of layers.
- It does not address loss balancing between tasks. The paper pairs it with a dynamic loss-weighting scheme.