Model techniques map
Techniquespost-trainingreinforcement learning algorithm

specific method · filed under post-training

Asynchronous Group Relative Policy Optimization

Group Relative Policy Optimization run fully asynchronously, including on very large batches.

sources
2
models
2
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1core 1

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

Fully asynchronous Group Relative Policy Optimization (GRPO) on very large batches

coreoptimizationin MiMo-V2.6-Flash-RLXiaomi

Fully asynchronous Group Relative Policy Optimization (GRPO) on very large batches

usedoptimizationin MiMo-V2.6-Pro-RLXiaomi

Filed alongside

Other methods under post-training :: reinforcement learning algorithm.