general family · filed under post-training
Generative Reward Model
A reward model that generates evaluations rather than returning only a scalar reward; the evidence also describes rubric-based response and trajectory scoring.
Also called GenRM, Generative Reward Model (GRM), Rubric-conditioned Generative Reward Model.
- sources
- 6
- models
- 4
- labs adopt it
- 2
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Further reading
Picked by hand, not extracted: where to read more, not evidence for anything on this page.
Evidence
7 spans quoted from the sources, strongest treatment first.
GenRM used for RLHF
We then develop detailed evaluation rubrics across multiple quality dimensions and employ a generative reward model to score responses based on these rubrics.
For general tasks, we employ a generative reward model where each prompt has its own rubrics for evaluation.
we curate rubric-guided RL data and employ a Generative Reward Model (GRM) to evaluate policy trajectories.
Nemotron 3 Ultra 550B-A55B GenRM: GenRM used for RLHF
Generative Reward Model replaces scalar reward models
Generative Reward Models (GenRM) (Wang2025HelpSteer3Preference) with reasoning capabilities help mitigate such reward hacking behaviors
Filed alongside
Other methods under post-training :: reward modelling.