specific method · filed under post-training
Abstention-aware reward for factual QA
Rewards answering factual questions when confidence is sufficient while allowing abstention or hedged guesses when it is not.
Also called Abstention-aware rewards for factual QA.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
The largest is short-form factual QA with abstention-aware rewards: answering only pays off when the model is likely to be right, so the optimal policy is to answer when confident and otherwise say “I don’t know” or give a hedged best guess.
usedpost trainingin InklingThinking Machines Lab
Filed alongside
Other methods under post-training :: reward modelling.
Generative Reward ModelGroupwise Agentic GradingGroupwise Reward SynthesisAdversarial screeningLength-adjusted RL rewardVerifier cross-checkingAgentic Generative Reward ModelBehavior rubricsBinary task verifierBinary terminal-verifier rewardCollaboration bonusDeterministic chain of checkersDirect RL optimization of a generative reward modelFive-dimension comparative grading of passing patchesHack-agent screeningHybrid reward systemLanguage consistency rewardMonitoring-only penalty strategyMulti-level reward formulationMultiplicative reward synthesisNegative checks for unintended side effectsOutcome Reward ModelPer-token tool-error reward shapingPrinciple-conditioned Generative Reward Model