specific method · filed under post-training
Instruction-following fine-tuning
Additional training intended to improve how accurately a model understands and follows user instructions.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Instruction following: Malibu 2.1 went through additional training to understand and follow user instructions more accurately, especially over the course of long conversations.
usedpost trainingin Malibu 2.1Poolside
Filed alongside
Other methods under post-training :: supervised fine-tuning.
Supervised fine-tuningFine-tuningLight SFT on teacher-distribution dataLoRA fine-tuningRejection samplingSelf-correction cold startSFT bootstrapping with synthetic dataAdversarial fine-tuningAutoregressive fine-tuningCold-started SFT modelEvaluation-based early stoppingFine-tuning with demographically balanced datasetsFrozen-network warmup for new special tokensInstruction hierarchy trainingLightweight speaker fine-tuningPrompt-based cold start for tool useReasoning-data system promptSafety training to an internal specificationSFT checkpoint for RL researchSFT on MiMo-generated task dataSpecialist-model trainingSupervised fine-tuning at 256k sequence lengthToken-budget-based SFT data blending