Model techniques map
Techniquespost-trainingpolicy distillation

specific method · filed under post-training

Multi-Prefix Multi-Teacher On-Policy Distillation

Multi-teacher on-policy distillation combining autonomous student rollouts with prefix-conditioned single-turn rollouts.

sources
3
models
2
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 3

Documented in

Evidence

3 spans quoted from the sources, strongest treatment first.

After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned single-turn rollouts (Teacher-Prefix and SFT-Prefix)

usedpost trainingin MiMo-V2.6-Pro-RLXiaomi

After mixed RL, we use Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2) to combine capabilities from teachers trained for different tasks

usedunclearin MiMo-V2.6Xiaomi

MOPD2 combines autonomous student rollouts with prefix-conditioned single-turn rollouts (Teacher-Prefix and SFT-Prefix)

usedpost trainingin MiMo-V2.6-Flash-RLXiaomi

Filed alongside

Other methods under post-training :: policy distillation.