implementation detail · filed under post-training
Multiplicative reward synthesis
Forms the training reward by multiplying the test reward by rubric scores.
- source
- 1
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
We synthesize the final training reward by multiplying this test reward by the two rubric scores
usedpost trainingin MiMo-V2.6-Pro and MiMo-V2.6-FlashXiaomi
Filed alongside
Other methods under post-training :: reward modelling.
Generative Reward ModelGroupwise Agentic GradingGroupwise Reward SynthesisAdversarial screeningLength-adjusted RL rewardVerifier cross-checkingAbstention-aware reward for factual QAAgentic Generative Reward ModelBehavior rubricsBinary task verifierBinary terminal-verifier rewardCollaboration bonusDeterministic chain of checkersDirect RL optimization of a generative reward modelFive-dimension comparative grading of passing patchesHack-agent screeningHybrid reward systemLanguage consistency rewardMonitoring-only penalty strategyMulti-level reward formulationNegative checks for unintended side effectsOutcome Reward ModelPer-token tool-error reward shapingPrinciple-conditioned Generative Reward Model