specific method · filed under optimization
KDA Context Parallelism
A context-parallel method that decomposes each segment's effect into a cumulative transition on the incoming state and a locally generated state.
Also called KDA Context Parallelism (KCP).
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
To preserve this dependence, we introduce KDA Context Parallelism (KCP), which decomposes the effect of each segment into two locally computable quantities, a cumulative transition acting on the incoming state and a state generated locally from zero.
usedunclearin Kimi K3Moonshot AI
Filed alongside
Other methods under optimization :: training parallelism.
Expert ParallelismTensor ParallelismMoonEPCommunication-Computation OverlapPipeline ParallelismAll-to-all Gradient Exchange with Local FP32 SummationCache-Based Pipeline CommunicationContext ParallelismData Replica Reduction over the Data Center NetworkData-Weighted Data ParallelismDynamic Context Parallelism for Large Multimodal SamplesFully Balanced Expert-Parallel TrainingGPU Planning Kernel for Redundant Expert MigrationHybrid ZeRO Bucket Assignment for MuonKnapsack-Based Balanced Assignment of Dense Parameter MatricesLoad-Balanced Image ShardingModified DualPipe 1F1B Pipeline Overlap for mHCNS-FLOP-Balanced Static Parameter PartitioningPipeline Payload ExtensionsPipeline ZeRO-2 Gradient Sharding with CPU Offloadingpipeline-bubble scheduling of ViT computationRedundant-Expert Capacity ReservationSConv-Aware Tensor-Parallel ShardingSequence Parallelism for Activations