Model techniques map
Techniquespost-trainingreward modelling

general family · filed under post-training

Reinforcement learning with verifiable rewards

Uses rewards that can be verified, but the evidence does not specify a particular verifier or reward mechanism.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

Use reinforcement learning with verifiable rewards to adapt NVIDIA Nemotron for specific domains and workflows

usedtraining objectivein NVIDIA NemotronNVIDIA

Filed alongside

Other methods under post-training :: reward modelling.