Model techniques map
Techniquesdata curationsynthetic data

general family · filed under data curation

Automated batch synthesis of RL training data

Synthesized-task pipelines automatically produce RL training data in batches with controllable difficulty and length.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

Through these pipelines, we can automatically and batch-produce RL training data that is correct, discriminative, and controllable in length and difficulty.

useddata curationin DeepSeek-V4.1 RL training dataDeepSeek

Filed alongside

Other methods under data curation :: synthetic data.