general family · filed under post-training
SFT followed by RL and on-policy distillation
A post-training sequence of supervised fine-tuning, reinforcement learning, and on-policy distillation.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- default
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
the recipe follows the standard paradigm of supervised fine-tuning (SFT) followed by reinforcement learning (RL) and on-policy distillation (OPD), without any modification beyond well-established practice used in DeepSeek-V4 development.
The overall recipe follows the standard paradigm of supervised fine-tuning (SFT) followed by reinforcement learning (RL) and on-policy distillation (OPD; Gu et al., 2024; Lu and Lab, 2025), without algorithmic modifications beyond well-established practices.
Filed alongside
Other methods under post-training.