specific method · filed under post-training
Freezing the MoE Router During Reinforcement Learning
Keeping the mixture-of-experts router fixed during RL training, with the cited evidence associating this choice with stable expert-load statistics.
Also called freeze the router for RL training, freeze the MoE router, Freezing the MoE router during RL, freezing the router.
- source
- 1
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
evaluated 1used 2
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
To keep training stable at scale, we freeze the MoE router
usedoptimizationin MiMo-V2.6Xiaomi
We therefore freeze the router for RL training; the frozen-router run keeps all three statistics flat
usedoptimizationin MiMo-V2.6-ProXiaomi
comparing runs with and without freezing the router
evaluatedunclearin MiMo-V2.6-ProXiaomi
Filed alongside
Other methods under post-training :: reinforcement learning algorithm.
Group Relative Policy OptimizationReinforcement LearningGroupwise Advantage RedistributionIcePopOff-Policy Sequence MaskingReinforcement Learning from Verifiable RewardsRollout Routing ReplayUnbiased KL EstimateAsynchronous Group Relative Policy OptimizationChain-of-Thought Reinforcement LearningDirect Double-Sided Importance SamplingKeep Sampling MaskMixed Reinforcement LearningPivot Reinforcement LearningReinforcement Learning Post-TrainingAbstention TrainingAdvantage ShapingAgentic RL Task MixCISPO with Length-Weighted Leave-One-Out Group-Relative AdvantagesConcatenated Routing ReplayDerived-Latency PenaltyDomain-Specialized RL ExpertsDomain-Specific GRPO TrainingDropping All-Zero-Advantage Groups