implementation detail · filed under data curation
Long-context data synthesis by permutation and concatenation
Multimodal documents and subtasks are permuted and concatenated to create long-context examples requiring information distributed across the context.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
To address this, we synthesize additional long-context data by carefully permuting and concatenating multimodal documents and sub-tasks, so that the embedded tasks can be solved only by attending to information scattered across the full 1M-token context.
usedunclearin Kimi K3Moonshot AI
Filed alongside
Other methods under data curation :: synthetic data.
Agentic data synthesis pipelineKnowledge distillation for synthetic dataAutomated data synthesis and environment-construction pipelinesAutomatic synthesis of task-oriented RL environmentsContainerized coding-environment construction with self-testing and trace removalCounterfactual data augmentationRecycling filtered documents into image-text pairsAutomated batch synthesis of RL training dataCoding environments from GitHub issue–PR pairsContainer buildability and verifiability checkCorpus rephrasing with fidelity verificationDynamic multi-agent synthetic-data generation loopEnd-to-end synthetic environment generationEnvironment, toolset, task, and solution synthesis pipelineFailure-case and negative-feedback-driven environment generationHeterogeneous answer-generation agentsIterative task-difficulty escalationKnowledge-graph-guided task synthesisLarge-scale environment synthesis and curationMatching synthesis-pipeline complexity to teacher capabilityMetadata-conditioned synthetic generationMocked tools for agent environmentsModular synthetic-data pipeline compositionMulti-agent collaborative environment construction