specific method · filed under post-training
Reinforcement Learning from Verifiable Rewards
Reinforcement learning using verifiable rewards, including a unified stage spanning multiple available environments.
Also called unified RLVR, RLVR.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 3
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
we conduct a unified RLVR (Reinforcement Learning with Verifiable Reward) training stage spanning all available environments
usedpost trainingin Nemotron 3 UltraNVIDIA
followed by unified RLVR over a wide mix of reasoning, agentic, code, safety, usability, and chat environments
usedpost trainingin Nemotron 3 UltraNVIDIA
RLVR (Reinforcement Learning with Verifiable Reward) training stage spanning all available environments
usedoptimizationin Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under post-training :: reinforcement learning algorithm.
Group Relative Policy OptimizationReinforcement LearningGroupwise Advantage RedistributionFreezing the MoE Router During Reinforcement LearningIcePopOff-Policy Sequence MaskingRollout Routing ReplayUnbiased KL EstimateAsynchronous Group Relative Policy OptimizationChain-of-Thought Reinforcement LearningDirect Double-Sided Importance SamplingKeep Sampling MaskMixed Reinforcement LearningPivot Reinforcement LearningReinforcement Learning Post-TrainingAbstention TrainingAdvantage ShapingAgentic RL Task MixCISPO with Length-Weighted Leave-One-Out Group-Relative AdvantagesConcatenated Routing ReplayDerived-Latency PenaltyDomain-Specialized RL ExpertsDomain-Specific GRPO TrainingDropping All-Zero-Advantage Groups