Model techniques map
Techniquespost-trainingreward modelling

specific method · filed under post-training

Deterministic chain of checkers

Applies deterministic reward checks in order, with the first failed check determining the reward.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

Rewards are produced by a deterministic chain of checkers applied to every terminated rollout in the following order, with the first failing check determining the reward.

corepost trainingin Laguna XS.2Poolside

Filed alongside

Other methods under post-training :: reward modelling.