general family · filed under post-training
Scalable reinforcement learning post-training
A scalable reinforcement-learning post-training framework, with the cited evidence not specifying a particular scaling mechanism.
Also called Scalable Reinforcement Learning Framework.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1core 1
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
By implementing a robust RL protocol and scaling post-training compute, DeepSeek-V3.2 performs comparably to GPT-5.
corepost trainingin DeepSeek-V3.2DeepSeek
A scalable reinforcement learning post-training framework further improves reasoning
usedpost trainingin DeepSeek-V3.2DeepSeek
Filed alongside
Other methods under post-training.
Scalable RL at Agent ScaleReinforcement learning for low-pass-rate tasksSFT followed by RL and on-policy distillationExtended post-trainingJoint SFT and RLPost-training on diverse domainsPost-training optimizationThree-stage post-training with multi-teacher on-policy distillationThree-stage Thinker post-training strategy