Model techniques map
Techniquespost-training

general family · filed under post-training

SFT followed by RL and on-policy distillation

A post-training sequence of supervised fine-tuning, reinforcement learning, and on-policy distillation.

source
1
model
1
lab adopt it
1
strongest
default

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1default 1

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

the recipe follows the standard paradigm of supervised fine-tuning (SFT) followed by reinforcement learning (RL) and on-policy distillation (OPD), without any modification beyond well-established practice used in DeepSeek-V4 development.

defaultpost trainingin DeepSeek-V4.1-FlashDeepSeek

The overall recipe follows the standard paradigm of supervised fine-tuning (SFT) followed by reinforcement learning (RL) and on-policy distillation (OPD; Gu et al., 2024; Lu and Lab, 2025), without algorithmic modifications beyond well-established practices.

usedunclearin DeepSeek-V4.1-FlashDeepSeek

Filed alongside

Other methods under post-training.