specific method · filed under post-training
Progressive Scaling of RL Data, Tasks, and Rollouts
Progressively increasing the data, tasks, and rollouts used in RL to extend capabilities across domains.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
progressively scale the data, tasks, and rollouts employed during RL, thereby extending the model's capabilities across textual, multimodal, and agentic domains
usedpost trainingin DeepSeek-V4.1-FlashDeepSeek
Filed alongside
Other methods under post-training :: reinforcement learning algorithm.
Group Relative Policy OptimizationReinforcement LearningGroupwise Advantage RedistributionFreezing the MoE Router During Reinforcement LearningIcePopOff-Policy Sequence MaskingReinforcement Learning from Verifiable RewardsRollout Routing ReplayUnbiased KL EstimateAsynchronous Group Relative Policy OptimizationChain-of-Thought Reinforcement LearningDirect Double-Sided Importance SamplingKeep Sampling MaskMixed Reinforcement LearningPivot Reinforcement LearningReinforcement Learning Post-TrainingAbstention TrainingAdvantage ShapingAgentic RL Task MixCISPO with Length-Weighted Leave-One-Out Group-Relative AdvantagesConcatenated Routing ReplayDerived-Latency PenaltyDomain-Specialized RL ExpertsDomain-Specific GRPO Training