Model techniques map
Taxonomypost-trainingreward modelling

taxonomy node · level 2

reward modelling

35 methods filed at this node or below it, from the sources of 14 models.

post-training :: reward modelling

Matching aids for the classifier: generative reward model; outcome reward model; rule-based verifiers; RLVR; hybrid reward system; test-difficulty driven code reward.

In this branch 35

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 35

Groupwise Agentic Grading core · 3 sources · 4 quotes
Groupwise Reward Synthesis core · 3 sources · 4 quotes
Binary task verifier core · 1 source · 1 quote
Binary terminal-verifier reward core · 1 source · 1 quote
Deterministic chain of checkers core · 1 source · 1 quote
Generative Reward Model used · 6 sources · 7 quotes
Adversarial screening used · 2 sources · 2 quotes
Length-adjusted RL reward used · 2 sources · 2 quotes
Verifier cross-checking used · 2 sources · 2 quotes
Abstention-aware reward for factual QA used · 1 source · 1 quote
Agentic Generative Reward Model used · 1 source · 1 quote
Behavior rubrics used · 1 source · 1 quote
Collaboration bonus used · 1 source · 1 quote
Hack-agent screening used · 1 source · 1 quote
Hybrid reward system used · 1 source · 1 quote
Language consistency reward used · 1 source · 1 quote
Monitoring-only penalty strategy used · 1 source · 1 quote
Multi-level reward formulation used · 1 source · 1 quote
Multiplicative reward synthesis used · 1 source · 1 quote
Outcome Reward Model used · 1 source · 1 quote
Per-token tool-error reward shaping used · 1 source · 1 quote
Public and hidden verifier pairing used · 1 source · 1 quote
Rule-based outcome reward used · 1 source · 1 quote
Rule-based verifier used · 1 source · 1 quote
Segment-level behavioral penalties used · 1 source · 1 quote
Solution rubrics used · 1 source · 1 quote
Test-difficulty-driven code reward used · 1 source · 1 quote
Training-time trajectory auditing used · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modeltechniques
DeepSeek-V4.1-Flash usedCollaboration bonus used—
NVIDIA-Nemotron-3-Ultra-550B-A55B usedGenerative Reward Model usedLength-adjusted RL reward usedPrinciple-conditioned Generative Reward Model used—
MiMo-V2.6-Flash coreGroupwise Agentic Grading coreGroupwise Reward Synthesis coreAdversarial screening usedBehavior rubrics usedFive-dimension comparative grading of passing patches usedHack-agent screening usedMonitoring-only penalty strategy usedMultiplicative reward synthesis usedNegative checks for unintended side effects usedRule-based or model-judged trajectory error detection usedSegment-level behavioral penalties usedSolution rubrics usedTraining-time trajectory auditing usedVerifier cross-checking used—
DeepSeek-V4-Flash usedDirect RL optimization of a generative reward model usedGenerative Reward Model used—
MiMo-V2.5 usedOutcome Reward Model usedRule-based verifier usedTest-difficulty-driven code reward used—
GLM-5.2 usedHybrid reward system usedMulti-level reward formulation used—
DeepSeek-V3.2 usedGenerative Reward Model usedLanguage consistency reward usedLength-adjusted RL reward usedRule-based outcome reward used—
DeepSeek-V4-Pro usedDirect RL optimization of a generative reward model usedGenerative Reward Model used—
Inkling usedAbstention-aware reward for factual QA used—
Kimi K3 usedAgentic Generative Reward Model usedPublic and hidden verifier pairing usedWeb-development reward with deterministic checks and model judging used—
Laguna-S-2.1 coreBinary task verifier coreBinary terminal-verifier reward coreDeterministic chain of checkers corePer-token tool-error reward shaping used—
MiMo-V2.5-Pro usedRule-based verifier usedTest-difficulty-driven code reward used—
MiMo-V2.6-Pro coreGroupwise Agentic Grading coreAdversarial screening usedBehavior rubrics usedFive-dimension comparative grading of passing patches usedGroupwise Reward Synthesis usedHack-agent screening usedMonitoring-only penalty strategy usedMultiplicative reward synthesis usedNegative checks for unintended side effects usedRule-based or model-judged trajectory error detection usedSegment-level behavioral penalties usedSolution rubrics usedTraining-time trajectory auditing usedVerifier cross-checking used—
NVIDIA-Nemotron-3.5-Lightning-30B-A3B usedReinforcement learning with verifiable rewards used—