specific method · filed under post-training
Groupwise Reward Synthesis
Builds task-specific rubrics from contrasting offline rollouts and combines rubric-based quality scores with test outcomes.
Also called Groupwise Reward Synthesis (GRS).
- sources
- 3
- models
- 2
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Evidence
4 spans quoted from the sources, strongest treatment first.
Groupwise Reward Synthesis (GRS) builds task-specific rubrics offline from contrasting rollouts and fuses rubric quality with test outcomes
Groupwise Reward Synthesis (GRS) builds task-specific rubrics offline from contrasting rollouts and fuses rubric quality with test outcomes
Through our proposed Groupwise Reward Synthesis (GRS) and Groupwise Advantage Redistribution (GAR), these fine-grained distinctions are converted into more informative learning signals
It compares multiple offline rollouts to construct task-specific rubrics, which are reused during training to score individual rollouts and combine their quality scores with test rewards.
Filed alongside
Other methods under post-training :: reward modelling.