general family · filed under data curation
Automated batch synthesis of RL training data
Synthesized-task pipelines automatically produce RL training data in batches with controllable difficulty and length.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Through these pipelines, we can automatically and batch-produce RL training data that is correct, discriminative, and controllable in length and difficulty.
useddata curationin DeepSeek-V4.1 RL training dataDeepSeek
Filed alongside
Other methods under data curation :: synthetic data.
Agentic data synthesis pipelineKnowledge distillation for synthetic dataAutomated data synthesis and environment-construction pipelinesAutomatic synthesis of task-oriented RL environmentsContainerized coding-environment construction with self-testing and trace removalCounterfactual data augmentationRecycling filtered documents into image-text pairsCoding environments from GitHub issue–PR pairsContainer buildability and verifiability checkCorpus rephrasing with fidelity verificationDynamic multi-agent synthetic-data generation loopEnd-to-end synthetic environment generationEnvironment, toolset, task, and solution synthesis pipelineFailure-case and negative-feedback-driven environment generationHeterogeneous answer-generation agentsIterative task-difficulty escalationKnowledge-graph-guided task synthesisLarge-scale environment synthesis and curationLong-context data synthesis by permutation and concatenationMatching synthesis-pipeline complexity to teacher capabilityMetadata-conditioned synthetic generationMocked tools for agent environmentsModular synthetic-data pipeline compositionMulti-agent collaborative environment construction