general family · filed under post-training
Model Distillation
A general reference to distillation, without evidence specifying a more particular mechanism.
Also called distillation.
- source
- 1
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
These progressive qualitative improvements align with the quantitative trajectory (61.7 → 64.0 → 72.4) observed in Table 6, jointly demonstrating the cumulative benefits of distillation from MiMo-V2.6 followed by RL
usedpost trainingin MiMo-V2.6-Distill-Qwen-9B
Filed alongside
Other methods under post-training :: policy distillation.
Multi-Teacher On-Policy DistillationOn-Policy DistillationPrefix-Conditioned On-Policy DistillationMulti-Prefix Multi-Teacher On-Policy DistillationSpecialist DistillationAutonomous Student RolloutsLarge-Scale On-Policy DistillationOn-Policy Cross-Stage DistillationAsynchronous Multi-Teacher On-Policy DistillationBehavior–Proximal Policy DecouplingDistillation Fine-Tuning on MiMo-Generated DataDistillation for Post-Training Data GenerationIcePop Token-Level Loss MaskingMulti-Objective Policy DistillationOff-Policy DistillationPer-Token On-Policy Distillation RewardSFT–RL–On-Policy Distillation PipelineTeacher-Trajectory and SFT-History Reuse