general family · filed under post-training
Chain-of-Thought Reinforcement Learning
Post-training with reinforcement-learning techniques described as similar to those used for OpenAI o3; the evidence does not specify a more precise algorithm.
Also called similar CoT RL techniques as OpenAI o3.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
After pre-training, we post-train the models using similar CoT RL techniques as OpenAI o3.
corepost trainingin gpt-oss-120b and gpt-oss-20bOpenAI
After pre-training, we post-train the models using similar CoT RL techniques as OpenAI o3.
corepost trainingin gpt-oss-120b and gpt-oss-20bOpenAI
Filed alongside
Other methods under post-training :: reinforcement learning algorithm.
Group Relative Policy OptimizationReinforcement LearningGroupwise Advantage RedistributionFreezing the MoE Router During Reinforcement LearningIcePopOff-Policy Sequence MaskingReinforcement Learning from Verifiable RewardsRollout Routing ReplayUnbiased KL EstimateAsynchronous Group Relative Policy OptimizationDirect Double-Sided Importance SamplingKeep Sampling MaskMixed Reinforcement LearningPivot Reinforcement LearningReinforcement Learning Post-TrainingAbstention TrainingAdvantage ShapingAgentic RL Task MixCISPO with Length-Weighted Leave-One-Out Group-Relative AdvantagesConcatenated Routing ReplayDerived-Latency PenaltyDomain-Specialized RL ExpertsDomain-Specific GRPO TrainingDropping All-Zero-Advantage Groups