specific method · filed under data curation
Counterfactual data augmentation
Data is augmented counterfactually as a mitigation strategy to align model behavior with desired behavior.
- sources
- 2
- models
- 2
- labs adopt it
- 0
- strongest
- mentioned
How sources treat it
One count per evidence span, weakest treatment to strongest.
mentioned 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
mitigation strategies like counterfactual data augmentation to align with the desired model behavior
mentioneddata curationNVIDIA
mitigation strategies like counterfactual data augmentation to align with the desired model behavior
mentioneddata curationin Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under data curation :: synthetic data.
Agentic data synthesis pipelineKnowledge distillation for synthetic dataAutomated data synthesis and environment-construction pipelinesAutomatic synthesis of task-oriented RL environmentsContainerized coding-environment construction with self-testing and trace removalRecycling filtered documents into image-text pairsAutomated batch synthesis of RL training dataCoding environments from GitHub issue–PR pairsContainer buildability and verifiability checkCorpus rephrasing with fidelity verificationDynamic multi-agent synthetic-data generation loopEnd-to-end synthetic environment generationEnvironment, toolset, task, and solution synthesis pipelineFailure-case and negative-feedback-driven environment generationHeterogeneous answer-generation agentsIterative task-difficulty escalationKnowledge-graph-guided task synthesisLarge-scale environment synthesis and curationLong-context data synthesis by permutation and concatenationMatching synthesis-pipeline complexity to teacher capabilityMetadata-conditioned synthetic generationMocked tools for agent environmentsModular synthetic-data pipeline compositionMulti-agent collaborative environment construction