Model techniques map
Techniquespost-trainingsupervised fine-tuning

general family · filed under post-training

Supervised fine-tuning

Supervised post-training used to establish or improve a model's behavior; the evidence does not specify one particular data or training recipe.

Also called SFT, Supervised Fine Tuning (SFT).

sources
14
models
7
labs adopt it
5
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 17

Documented in

Further reading

Picked by hand, not extracted: where to read more, not evidence for anything on this page.

Evidence

17 spans quoted from the sources, strongest treatment first.

Through joint optimization of SFT and RL

usedpost trainingin Hy3Tencent Hunyuan

The SFT stage establishes a high-quality cold-start policy for the subsequent RL stage.

usedunclearin Kimi K3Moonshot AI

Through joint optimization of SFT and RL

usedtraining objectivein Hy3Tencent Hunyuan

post-trained using Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation

usedpost trainingin Nemotron 3 UltraNVIDIA

To train the student model, we start from the pre-trained base model and perform Supervised Fine-Tuning (SFT) in two stages

usedpost trainingin Nemotron 3 UltraNVIDIA

Supervised Fine-Tuning (SFT) to build strong, foundational instruction-following skills

usedpost trainingin MiMo-V2.5-ProXiaomi

Post-trained with enhanced pipeline involving Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD) for improved model accuracy.

usedpost trainingin Nemotron 3 UltraNVIDIA

each specialist goes through its own Supervised Fine-Tuning (SFT) on domain-specific data

usedpost trainingin DeepSeek-V4DeepSeek

The model was further fine-tuned on synthetic code, math, science, tool calling, instruction following, structured outputs, and general knowledge data.

usedpost trainingin Nemotron 3.5 LightningNVIDIA

Post-training incorporates SFT, large-scale agentic RL, and Multi-Teacher On-Policy Distillation (MOPD)

usedpost trainingin MiMo-V2.5Xiaomi

Supervised Fine-Tuning (SFT) in two stages

usedpost trainingin Nemotron 3 UltraNVIDIA

To improve upon SFT model, we conduct a unified RLVR

usedpost trainingin Nemotron 3 UltraNVIDIA

post-trained using SFT, RL, and MOPD

usedpost trainingin Nemotron 3 UltraNVIDIA

independent cultivation of domain-specific experts (through SFT and RL with GRPO)

usedpost trainingin DeepSeek-V4DeepSeek

Post-training incorporates SFT, large-scale agentic RL, and Multi-Teacher On-Policy Distillation (MOPD)

usedpost trainingin MiMo-V2.5Xiaomi

general Supervised Fine-Tuning (SFT)

usedpost trainingin MiMo-V2-FlashXiaomi

supervised fine-tuning and agentic post-training

usedtraining objectivein MiMo-V2.5Xiaomi

Filed alongside

Other methods under post-training :: supervised fine-tuning.