Model techniques map
Techniquespost-training

specific method · filed under post-training

Reinforcement learning for low-pass-rate tasks

A two-stage post-training approach that reserves reinforcement learning for tasks the model cannot yet solve at a high pass rate.

Also called RL reserved for tasks with low pass rate.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1core 1

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

The majority of what separates S 2.1 from the XS models comes from post-training, in two stages: an SFT stage that bootstraps capabilities partly with synthetic data, then RL, reserved for tasks the model can't yet solve at a high pass rate.

corepost trainingin Laguna S 2.1Poolside

then RL, reserved for tasks the model can't yet solve at a high pass rate.

usedpost trainingin Laguna S 2.1Poolside

Filed alongside

Other methods under post-training.