specific method · filed under post-training
SFT on MiMo-generated task data
Supervised fine-tuning of Qwen3.5-9B on MiMo-generated data spanning coding, general-domain, visual, and cybersecurity tasks.
- source
- 1
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
We obtain the model by supervised fine-tuning Qwen3.5-9B (Qwen Team, 2025) on MiMo-generated data spanning coding, general-domain, visual, and cybersecurity tasks.
usedpost trainingin MiMo-V2.6-Distill-Qwen-9BXiaomi
Filed alongside
Other methods under post-training :: supervised fine-tuning.
Supervised fine-tuningFine-tuningLight SFT on teacher-distribution dataLoRA fine-tuningRejection samplingSelf-correction cold startSFT bootstrapping with synthetic dataAdversarial fine-tuningAutoregressive fine-tuningCold-started SFT modelEvaluation-based early stoppingFine-tuning with demographically balanced datasetsFrozen-network warmup for new special tokensInstruction hierarchy trainingInstruction-following fine-tuningLightweight speaker fine-tuningPrompt-based cold start for tool useReasoning-data system promptSafety training to an internal specificationSFT checkpoint for RL researchSpecialist-model trainingSupervised fine-tuning at 256k sequence lengthToken-budget-based SFT data blending