implementation detail · filed under post-training
Per-token tool-error reward shaping
Assigns shaping feedback at the failing token or step to improve credit assignment for malformed tool calls.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
This per-token shaping focuses credit assignment on the failing step and discourages long chains of malformed tool calls before falling back on the terminal verifier.
usedpost trainingin Laguna XS.2Poolside
Filed alongside
Other methods under post-training :: reward modelling.
Generative Reward ModelGroupwise Agentic GradingGroupwise Reward SynthesisAdversarial screeningLength-adjusted RL rewardVerifier cross-checkingAbstention-aware reward for factual QAAgentic Generative Reward ModelBehavior rubricsBinary task verifierBinary terminal-verifier rewardCollaboration bonusDeterministic chain of checkersDirect RL optimization of a generative reward modelFive-dimension comparative grading of passing patchesHack-agent screeningHybrid reward systemLanguage consistency rewardMonitoring-only penalty strategyMulti-level reward formulationMultiplicative reward synthesisNegative checks for unintended side effectsOutcome Reward ModelPrinciple-conditioned Generative Reward Model