Model techniques map
Taxonomypost-trainingsupervised fine-tuning

taxonomy node · level 2

supervised fine-tuning

24 methods filed at this node or below it, from the sources of 17 models.

post-training :: supervised fine-tuning

Matching aids for the classifier: SFT; cold-started SFT model; specialist training.

In this branch 24

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 24

LoRA fine-tuning core · 2 sources · 2 quotes
Light SFT on teacher-distribution data core · 1 source · 2 quotes
SFT checkpoint for RL research core · 1 source · 1 quote
Supervised fine-tuning used · 14 sources · 17 quotes
Rejection sampling used · 2 sources · 2 quotes
Self-correction cold start used · 2 sources · 2 quotes
SFT bootstrapping with synthetic data used · 2 sources · 2 quotes
Adversarial fine-tuning used · 1 source · 1 quote
Autoregressive fine-tuning used · 1 source · 1 quote
Cold-started SFT model used · 1 source · 1 quote
Evaluation-based early stopping used · 1 source · 1 quote
Instruction hierarchy training used · 1 source · 1 quote
Instruction-following fine-tuning used · 1 source · 1 quote
Lightweight speaker fine-tuning used · 1 source · 1 quote
Prompt-based cold start for tool use used · 1 source · 1 quote
Reasoning-data system prompt used · 1 source · 1 quote
SFT on MiMo-generated task data used · 1 source · 1 quote
Specialist-model training used · 1 source · 1 quote
Token-budget-based SFT data blending used · 1 source · 1 quote
Fine-tuning optional · 3 sources · 3 quotes

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modeltechniques
DeepSeek-V4.1-Flash usedAutoregressive fine-tuning used—
NVIDIA-Nemotron-3-Ultra-550B-A55B coreLight SFT on teacher-distribution data coreRejection sampling usedSupervised fine-tuning usedToken-budget-based SFT data blending used—
MiMo-V2.6-Flash coreSFT checkpoint for RL research coreSelf-correction cold start usedSFT on MiMo-generated task data used—
MiMo-V2.5 usedCold-started SFT model usedSupervised fine-tuning used—
Hy3 usedSupervised fine-tuning usedFine-tuning optional—
GLM-5.2 usedRejection sampling used—
DeepSeek-V3.2 usedPrompt-based cold start for tool use usedReasoning-data system prompt used—
DeepSeek-V4-Pro usedSpecialist-model training usedSupervised fine-tuning used—
Gemma 4 31B mentionedFine-tuning mentioned—
Inkling coreLoRA fine-tuning coreSafety training to an internal specification usedSFT bootstrapping with synthetic data used—
Kimi K3 usedSupervised fine-tuning used—
Laguna-S-2.1 usedEvaluation-based early stopping usedFrozen-network warmup for new special tokens usedInstruction-following fine-tuning usedSFT bootstrapping with synthetic data used—
MiMo-V2.5-Pro usedCold-started SFT model usedSupervised fine-tuning used—
MiMo-V2.6-Pro coreSFT checkpoint for RL research coreSelf-correction cold start usedSFT on MiMo-generated task data used—
NVIDIA-Nemotron-3.5-Lightning-30B-A3B usedLoRA fine-tuning usedSupervised fine-tuning usedSupervised fine-tuning at 256k sequence length usedFine-tuning with demographically balanced datasets mentioned—
Qwen3.5-397B-A17B usedLightweight speaker fine-tuning used—
gpt-oss-120b usedAdversarial fine-tuning usedInstruction hierarchy training usedFine-tuning optional—