specific method · filed under data curation
Training-task generation from real commit history
Software-engineering training tasks are grounded in reproductions of real commits and repository history.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Software-engineering tasks are mostly grounded in real code history: the largest source reproduces real commits (~38,000 tasks across ~17,000 repositories).
useddata curationin Laguna S 2.1Poolside
Filed alongside
Other methods under data curation :: synthetic data.
Agentic data synthesis pipelineKnowledge distillation for synthetic dataAutomated data synthesis and environment-construction pipelinesAutomatic synthesis of task-oriented RL environmentsContainerized coding-environment construction with self-testing and trace removalCounterfactual data augmentationRecycling filtered documents into image-text pairsAutomated batch synthesis of RL training dataCoding environments from GitHub issue–PR pairsContainer buildability and verifiability checkCorpus rephrasing with fidelity verificationDynamic multi-agent synthetic-data generation loopEnd-to-end synthetic environment generationEnvironment, toolset, task, and solution synthesis pipelineFailure-case and negative-feedback-driven environment generationHeterogeneous answer-generation agentsIterative task-difficulty escalationKnowledge-graph-guided task synthesisLarge-scale environment synthesis and curationLong-context data synthesis by permutation and concatenationMatching synthesis-pipeline complexity to teacher capabilityMetadata-conditioned synthetic generationMocked tools for agent environmentsModular synthetic-data pipeline composition