specific method · filed under post-training
Multi-Teacher On-Policy Distillation
On-policy distillation that consolidates capabilities from multiple teachers using states induced by the student.
Also called Mixture of On-Policy Distillation, MOPD, MOPD post-training paradigm, Multi-Domain On-Policy Distillation, Multi-Domain On-Policy Distillation (MOPD), multi-teacher On-Policy Distillation (OPD).
- sources
- 12
- models
- 6
- labs adopt it
- 4
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Further reading
Picked by hand, not extracted: where to read more, not evidence for anything on this page.
Evidence
17 spans quoted from the sources, strongest treatment first.
We adopt Multi-Teacher On-Policy Distillation (MOPD) to consolidate these domain-specialized capabilities across varying reasoning efforts into a unified model
we augment the pipeline with Multi-teacher On-Policy Distillation (MOPD) (Yang et al., 2026; Lu & Lab, 2025; Xiao et al., 2026), enabling both broad capability acquisition and targeted specialization.
MOPD provides a dense token-level learning signal from the relevant teacher distribution.
we employ multi-teacher On-Policy Distillation (OPD) as the primary technique for merging expert capabilities into the final model.
Multi-Teacher On-Policy Distillation (MOPD)
Multi-teacher On-Policy Distillation (MOPD) consolidated these teachers into Ultra
MiMo-V2-Flash introduces a novel Multi-Teacher On-Policy Distillation (MOPD) paradigm
through its hybrid Sliding Window Attention architecture, lightweight Multi-Token Prediction, and the MOPD post-training paradigm.
Domain- and effort-specialized policies are consolidated into a unified model through multi-teacher on-policy distillation [75, 134, 29].
Multi-teacher On-Policy Distillation (MOPD) consolidated these teachers into Ultra
The model underwent Multi-Domain On-Policy Distillation (MOPD) to improve reasoning across many task types while staying efficient
Post-trained with enhanced pipeline involving Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD) for improved model accuracy.
Multi-Domain On-Policy Distillation (MOPD) to improve reasoning across many task types while staying efficient
Post-training incorporates SFT, large-scale agentic RL, and Multi-Teacher On-Policy Distillation (MOPD)
MOPD trains the student to match the corresponding teacher on states induced by the student itself
Post-training incorporates SFT, large-scale agentic RL, and Multi-Teacher On-Policy Distillation (MOPD)
Filed alongside
Other methods under post-training :: policy distillation.