Model techniques map
Techniquesdata curationsynthetic data

general family · filed under data curation

Automated data synthesis and environment-construction pipelines

A broad family of large-scale automated pipelines for synthesizing data and constructing environments.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 2

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

We develop large-scale automated pipelines for data synthesis and environment construction, and progressively scale the data, tasks, and rollouts employed during RL.

coredata curationin DeepSeek-V4.1-FlashDeepSeek

our efforts are concentrated almost entirely on what the model is trained on rather than how it is optimized: we invest in large-scale, automated pipelines for data synthesis and environment construction.

coreunclearin DeepSeek-V4.1-FlashDeepSeek

Filed alongside

Other methods under data curation :: synthetic data.