Model techniques map
Techniquespost-trainingpolicy distillation

specific method · filed under post-training

Multi-Teacher On-Policy Distillation

On-policy distillation that consolidates capabilities from multiple teachers using states induced by the student.

Also called Mixture of On-Policy Distillation, MOPD, MOPD post-training paradigm, Multi-Domain On-Policy Distillation, Multi-Domain On-Policy Distillation (MOPD), multi-teacher On-Policy Distillation (OPD).

sources
12
models
6
labs adopt it
4
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 9core 8

Documented in

Further reading

Picked by hand, not extracted: where to read more, not evidence for anything on this page.

Evidence

17 spans quoted from the sources, strongest treatment first.

We adopt Multi-Teacher On-Policy Distillation (MOPD) to consolidate these domain-specialized capabilities across varying reasoning efforts into a unified model

coreunclearin Kimi K3Moonshot AI

we augment the pipeline with Multi-teacher On-Policy Distillation (MOPD) (Yang et al., 2026; Lu & Lab, 2025; Xiao et al., 2026), enabling both broad capability acquisition and targeted specialization.

corepost trainingin Nemotron 3 UltraNVIDIA

MOPD provides a dense token-level learning signal from the relevant teacher distribution.

coreunclearin Nemotron 3 UltraNVIDIA

we employ multi-teacher On-Policy Distillation (OPD) as the primary technique for merging expert capabilities into the final model.

corepost trainingin DeepSeek-V4DeepSeek

Multi-Teacher On-Policy Distillation (MOPD)

corepost trainingin MiMo-V2.5-ProXiaomi

Multi-teacher On-Policy Distillation (MOPD) consolidated these teachers into Ultra

corepost trainingin Nemotron 3 UltraNVIDIA

MiMo-V2-Flash introduces a novel Multi-Teacher On-Policy Distillation (MOPD) paradigm

corepost trainingin MiMo-V2-FlashXiaomi

through its hybrid Sliding Window Attention architecture, lightweight Multi-Token Prediction, and the MOPD post-training paradigm.

corepost trainingin MiMo-V2-FlashXiaomi

Domain- and effort-specialized policies are consolidated into a unified model through multi-teacher on-policy distillation [75, 134, 29].

usedpost trainingin Kimi K3Moonshot AI

Multi-teacher On-Policy Distillation (MOPD) consolidated these teachers into Ultra

usedpost trainingin Nemotron 3 UltraNVIDIA

The model underwent Multi-Domain On-Policy Distillation (MOPD) to improve reasoning across many task types while staying efficient

usedoptimizationin Nemotron 3 UltraNVIDIA

Post-trained with enhanced pipeline involving Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD) for improved model accuracy.

usedpost trainingin Nemotron 3 UltraNVIDIA

Multi-Domain On-Policy Distillation (MOPD) to improve reasoning across many task types while staying efficient

usedtraining objectivein Nemotron 3 UltraNVIDIA

Post-training incorporates SFT, large-scale agentic RL, and Multi-Teacher On-Policy Distillation (MOPD)

usedpost trainingin MiMo-V2.5Xiaomi

MOPD trains the student to match the corresponding teacher on states induced by the student itself

usedoptimizationin Nemotron 3 UltraNVIDIA

Post-training incorporates SFT, large-scale agentic RL, and Multi-Teacher On-Policy Distillation (MOPD)

usedpost trainingin MiMo-V2.5Xiaomi

RL and MOPD

usedoptimizationin MiMo-V2.5Xiaomi

Filed alongside

Other methods under post-training :: policy distillation.