implementation detail · filed under post-training
Five-dimension comparative grading of passing patches
Compares passing code patches on solution suitability, implementation precision, minimality, unintended effects, and craftsmanship.
- source
- 1
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
comparing passing patches along five dimensions: the suitability of the solution approach, precision in implementing that approach without omissions or unnecessary fallbacks, minimality relative to the necessary changes, avoidance of unintended effects outside the task, and craftsmanship consistent with codebase conventions.
usedpost trainingin MiMo-V2.6-Pro and MiMo-V2.6-FlashXiaomi
Filed alongside
Other methods under post-training :: reward modelling.
Generative Reward ModelGroupwise Agentic GradingGroupwise Reward SynthesisAdversarial screeningLength-adjusted RL rewardVerifier cross-checkingAbstention-aware reward for factual QAAgentic Generative Reward ModelBehavior rubricsBinary task verifierBinary terminal-verifier rewardCollaboration bonusDeterministic chain of checkersDirect RL optimization of a generative reward modelHack-agent screeningHybrid reward systemLanguage consistency rewardMonitoring-only penalty strategyMulti-level reward formulationMultiplicative reward synthesisNegative checks for unintended side effectsOutcome Reward ModelPer-token tool-error reward shapingPrinciple-conditioned Generative Reward Model