Model techniques map
Techniquespost-trainingsupervised fine-tuning

specific method · filed under post-training

SFT checkpoint for RL research

A supervised fine-tuning checkpoint presented as a starting point for reinforcement-learning research rather than as a finished assistant.

Also called SFT starting point for RL research.

source
1
models
2
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

this checkpoint is an SFT starting point for RL research, not a finished assistant

coreunclearin MiMo-V2.6-Distill-Qwen-9BXiaomi

Filed alongside

Other methods under post-training :: supervised fine-tuning.