Model techniques map
Taxonomydata curationsynthetic data

taxonomy node · level 2

synthetic data

56 methods filed at this node or below it, from the sources of 13 models.

data curation :: synthetic data

Matching aids for the classifier: synthetic data generation; agentic data synthesis pipeline; counterfactual data augmentation; task-specific data injection.

In this branch 56

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 56

Agentic data synthesis pipeline core · 4 sources · 4 quotes
Knowledge distillation for synthetic data used · 3 sources · 3 quotes
Heterogeneous answer-generation agents used · 1 source · 1 quote
Iterative task-difficulty escalation used · 1 source · 1 quote
Knowledge-graph-guided task synthesis used · 1 source · 1 quote
Metadata-conditioned synthetic generation used · 1 source · 1 quote
Mocked tools for agent environments used · 1 source · 1 quote
Multi-agent task-attempt validation used · 1 source · 1 quote
Programmatic multimodal data generation used · 1 source · 1 quote
Reward-driven task-construction training used · 1 source · 1 quote
Rubric-guided RL data used · 1 source · 1 quote
Search-based question construction used · 1 source · 1 quote
Search-capable answer verification agent used · 1 source · 1 quote
Synthetic data generation used · 1 source · 1 quote
Synthetic data generation and augmentation used · 1 source · 1 quote
Synthetic privacy-preserving personas used · 1 source · 1 quote
Synthetic system-message augmentation used · 1 source · 1 quote
Targeted synthetic data rephrasing used · 1 source · 1 quote
Tool-set and schema randomization used · 1 source · 1 quote
Counterfactual data augmentation mentioned · 2 sources · 2 quotes

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modeltechniques
DeepSeek-V4.1-Flash coreAutomated data synthesis and environment-construction pipelines coreProcedural construction and scaling of interactive agent environments coreProgressively scaled synthesis of agent tasks and environments coreSynthesis of verifiable tasks with solutions and reward signals coreAutomated batch synthesis of RL training data usedContainer buildability and verifiability check usedContainerized coding-environment construction with self-testing and trace removal usedFailure-case and negative-feedback-driven environment generation usedMocked tools for agent environments usedMulti-agent collaborative environment construction usedMulti-agent task-attempt validation usedRecycling filtered documents into image-text pairs usedReward-driven task-construction training usedTask formalization as a problem–environment–verification triplet used—
Hy4-preview usedTraining-data construction around expert work products used—
NVIDIA-Nemotron-3-Ultra-550B-A55B usedKnowledge distillation for synthetic data usedSynthetic data generation usedCounterfactual data augmentation mentioned—
MiMo-V2.6-Flash usedLarge-scale environment synthesis and curation usedMulti-agent synthesis of local software mocks usedSFT teachers trained on synthetic demonstrations used—
GLM-5.3 usedEnd-to-end synthetic environment generation usedVerifier synthesis without reference solutions usedVulnerability-focused training data and environments used—
GLM-5.2 usedAgentic data synthesis pipeline used—
DeepSeek-V3.2 coreAgentic data synthesis pipeline coreAutomatic synthesis of task-oriented RL environments coreCoding environments from GitHub issue–PR pairs usedEnvironment, toolset, task, and solution synthesis pipeline usedHeterogeneous answer-generation agents usedIterative task-difficulty escalation usedMulti-agent search-agent training-data generation usedSearch-based question construction usedSearch-capable answer verification agent used—
DeepSeek-V4-Pro usedRubric-guided RL data usedTask-specific data injection during mid-training used—
Inkling usedSynthetic data generation and augmentation usedTool-set and schema randomization used—
Kimi K3 usedCorpus rephrasing with fidelity verification usedKnowledge-graph-guided task synthesis usedLong-context data synthesis by permutation and concatenation usedMulti-stage verification with human-in-the-loop annotation usedProgrammatic multimodal data generation usedSandboxed visual reasoning with a Python interpreter usedSynthetic trajectory generation with domain-specialized models used—
Laguna-S-2.1 coreDynamic multi-agent synthetic-data generation loop coreModular synthetic-data pipeline composition coreMatching synthesis-pipeline complexity to teacher capability usedMetadata-conditioned synthetic generation usedMulti-stage synthetic-data generation cascade usedSynthetic data augmentation of organic data usedSynthetic system-message augmentation usedTargeted synthetic data rephrasing usedTraining-task generation from real commit history usedTwo-sided test validation for synthetic coding tasks used—
MiMo-V2.6-Pro usedLarge-scale environment synthesis and curation usedMulti-agent synthesis of local software mocks usedSFT teachers trained on synthetic demonstrations used—
NVIDIA-Nemotron-3.5-Lightning-30B-A3B usedKnowledge distillation for synthetic data usedSynthetic data generation, filtering, and curation usedSynthetic privacy-preserving personas usedCounterfactual data augmentation mentioned—