specific method · filed under data curation
Corpus rephrasing with fidelity verification
Knowledge and mathematics corpora are rephrased with varied styles and perspectives, then checked against source documents for fidelity.
Also called rephrasing recipe.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
we rephrase knowledge and mathematics corpora with style and perspective-diverse prompting, chunk-wise autoregressive generation, and fidelity verification against the source documents.
useddata curationin Kimi K3Moonshot AI
Filed alongside
Other methods under data curation :: synthetic data.
Agentic data synthesis pipelineKnowledge distillation for synthetic dataAutomated data synthesis and environment-construction pipelinesAutomatic synthesis of task-oriented RL environmentsContainerized coding-environment construction with self-testing and trace removalCounterfactual data augmentationRecycling filtered documents into image-text pairsAutomated batch synthesis of RL training dataCoding environments from GitHub issue–PR pairsContainer buildability and verifiability checkDynamic multi-agent synthetic-data generation loopEnd-to-end synthetic environment generationEnvironment, toolset, task, and solution synthesis pipelineFailure-case and negative-feedback-driven environment generationHeterogeneous answer-generation agentsIterative task-difficulty escalationKnowledge-graph-guided task synthesisLarge-scale environment synthesis and curationLong-context data synthesis by permutation and concatenationMatching synthesis-pipeline complexity to teacher capabilityMetadata-conditioned synthetic generationMocked tools for agent environmentsModular synthetic-data pipeline compositionMulti-agent collaborative environment construction