specific method · filed under post-training
Light SFT on teacher-distribution data
A light supervised fine-tuning stage using data drawn from a teacher's training distribution.
Also called teacher-distribution SFT warmup.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1core 1
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
the student undergoes a very light SFT on data drawn from the teacher’s training distribution.
coreunclearin Nemotron 3 UltraNVIDIA
before pivot RL, we performed light SFT directly on the student Ultra model.
usedunclearin Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under post-training :: supervised fine-tuning.
Supervised fine-tuningFine-tuningLoRA fine-tuningRejection samplingSelf-correction cold startSFT bootstrapping with synthetic dataAdversarial fine-tuningAutoregressive fine-tuningCold-started SFT modelEvaluation-based early stoppingFine-tuning with demographically balanced datasetsFrozen-network warmup for new special tokensInstruction hierarchy trainingInstruction-following fine-tuningLightweight speaker fine-tuningPrompt-based cold start for tool useReasoning-data system promptSafety training to an internal specificationSFT checkpoint for RL researchSFT on MiMo-generated task dataSpecialist-model trainingSupervised fine-tuning at 256k sequence lengthToken-budget-based SFT data blending