Model techniques map
Techniquespost-trainingreinforcement learning algorithm

general family · filed under post-training

Reinforcement Learning Post-Training

Post-training with RL; the evidence does not identify a particular algorithm.

Also called RL post-training, larger-scale RL post-training.

sources
2
models
2
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

mentioned 1used 1

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

RL Post-training

usedpost trainingin Hy3Tencent Hunyuan

New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models

mentionedpost trainingin DeepSeek-V4.1-FlashDeepSeek

Filed alongside

Other methods under post-training :: reinforcement learning algorithm.