general family · filed under post-training
Multi-harness training
Jointly optimizing a model across multiple harnesses, treating harness diversity as a training dimension.
Also called Multi-harness reinforcement learning.
- source
- 1
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
We therefore treat harness diversity as an additional training dimension alongside diversity in tasks and environments.
we jointly optimize the model across four mini-harnesses and evaluate it on these harnesses and three additional held-out harnesses.
we jointly optimize the model across four mini-harnesses and evaluate it on these harnesses and three additional held-out harnesses.
Filed alongside
Other methods under post-training :: agentic post-training.