specific method · filed under post-training
Groupwise Advantage Redistribution
An online method that ranks passing trajectories within a group and redistributes positive advantage toward higher-quality solutions.
Also called Groupwise Advantage Redistribution (GAR).
- sources
- 3
- models
- 2
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Evidence
4 spans quoted from the sources, strongest treatment first.
Groupwise Advantage Redistribution (GAR) ranks passing trajectories online and moves advantage toward higher-quality solutions.
Groupwise Advantage Redistribution (GAR) ranks passing trajectories online and moves advantage toward higher-quality solutions.
Through our proposed Groupwise Reward Synthesis (GRS) and Groupwise Advantage Redistribution (GAR), these fine-grained distinctions are converted into more informative learning signals
Its online agentic grader jointly examines successful and failed trajectories within each mixed-outcome group, ranks passing solutions, and redistributes positive advantage toward higher-quality passing trajectories.
Filed alongside
Other methods under post-training :: reinforcement learning algorithm.