specific method · filed under post-training
Mid-training alignment data
Alignment examples derived from reward-hacking cases and used during mid-training.
- source
- 1
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
For each case, MiMo reflects on the faulty reasoning, revises the relevant turn, and continues with actions grounded in the task specification.
usedpost trainingin MiMo-V2.6-Pro and MiMo-V2.6-FlashXiaomi
Filed alongside
Other methods under post-training :: mid-training & continual pretraining.