specific method · filed under post-training
Length-adjusted RL reward
Adjusts reinforcement-learning rewards based on response length; the evidence does not specify a particular adjustment formula.
Also called length penalty reward, length penalty.
- sources
- 2
- models
- 2
- labs adopt it
- 2
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
length-based adjustments are applied to the RL rewards for them.
usedunclearin Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under post-training :: reward modelling.
Generative Reward ModelGroupwise Agentic GradingGroupwise Reward SynthesisAdversarial screeningVerifier cross-checkingAbstention-aware reward for factual QAAgentic Generative Reward ModelBehavior rubricsBinary task verifierBinary terminal-verifier rewardCollaboration bonusDeterministic chain of checkersDirect RL optimization of a generative reward modelFive-dimension comparative grading of passing patchesHack-agent screeningHybrid reward systemLanguage consistency rewardMonitoring-only penalty strategyMulti-level reward formulationMultiplicative reward synthesisNegative checks for unintended side effectsOutcome Reward ModelPer-token tool-error reward shapingPrinciple-conditioned Generative Reward Model