general family · filed under post-training
Groupwise Agentic Grading
Uses an agentic grader to compare rollouts within a group and provide reward signals beyond binary test outcomes.
- sources
- 3
- models
- 2
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2core 2
Documented in
Evidence
4 spans quoted from the sources, strongest treatment first.
we introduce groupwise agentic grading to provide more informative reward signals beyond binary test cases
coreevaluation onlyin MiMo-V2.6 RL trainingXiaomi
An agentic grader compares rollouts within each group
coretraining objectivein MiMo-V2.6-Flash-RLXiaomi
An agentic grader compares rollouts within each group
usedtraining objectivein MiMo-V2.6-Pro-RLXiaomi
the compute devoted to groupwise agentic grading.
usedtraining objectivein MiMo-V2.6
Filed alongside
Other methods under post-training :: reward modelling.
Generative Reward ModelGroupwise Reward SynthesisAdversarial screeningLength-adjusted RL rewardVerifier cross-checkingAbstention-aware reward for factual QAAgentic Generative Reward ModelBehavior rubricsBinary task verifierBinary terminal-verifier rewardCollaboration bonusDeterministic chain of checkersDirect RL optimization of a generative reward modelFive-dimension comparative grading of passing patchesHack-agent screeningHybrid reward systemLanguage consistency rewardMonitoring-only penalty strategyMulti-level reward formulationMultiplicative reward synthesisNegative checks for unintended side effectsOutcome Reward ModelPer-token tool-error reward shapingPrinciple-conditioned Generative Reward Model