specific method · filed under data curation
Metadata-conditioned synthetic generation
Known output or process metadata is supplied to the generator to reduce the generation challenge.
Also called Help LLMs through metadata.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
We hand our generator chain G anything we already know about the output or how to achieve it through supplying metadata, reducing the challenge on the generator alone.
useddata curationin Laguna XS.2Poolside
Filed alongside
Other methods under data curation :: synthetic data.
Agentic data synthesis pipelineKnowledge distillation for synthetic dataAutomated data synthesis and environment-construction pipelinesAutomatic synthesis of task-oriented RL environmentsContainerized coding-environment construction with self-testing and trace removalCounterfactual data augmentationRecycling filtered documents into image-text pairsAutomated batch synthesis of RL training dataCoding environments from GitHub issue–PR pairsContainer buildability and verifiability checkCorpus rephrasing with fidelity verificationDynamic multi-agent synthetic-data generation loopEnd-to-end synthetic environment generationEnvironment, toolset, task, and solution synthesis pipelineFailure-case and negative-feedback-driven environment generationHeterogeneous answer-generation agentsIterative task-difficulty escalationKnowledge-graph-guided task synthesisLarge-scale environment synthesis and curationLong-context data synthesis by permutation and concatenationMatching synthesis-pipeline complexity to teacher capabilityMocked tools for agent environmentsModular synthetic-data pipeline compositionMulti-agent collaborative environment construction