implementation detail · filed under optimization
pipeline-bubble scheduling of ViT computation
Also called scheduled into pipeline bubbles.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
the remaining forward passes are scheduled into pipeline bubbles, and the backward passes are handled analogously.
usedsoftware implementationin Kimi K3Moonshot AI
Filed alongside
Other methods under optimization :: training parallelism.
Expert ParallelismTensor ParallelismMoonEPCommunication-Computation OverlapPipeline ParallelismAll-to-all Gradient Exchange with Local FP32 SummationCache-Based Pipeline CommunicationContext ParallelismData Replica Reduction over the Data Center NetworkData-Weighted Data ParallelismDynamic Context Parallelism for Large Multimodal SamplesFully Balanced Expert-Parallel TrainingGPU Planning Kernel for Redundant Expert MigrationHybrid ZeRO Bucket Assignment for MuonKDA Context ParallelismKnapsack-Based Balanced Assignment of Dense Parameter MatricesLoad-Balanced Image ShardingModified DualPipe 1F1B Pipeline Overlap for mHCNS-FLOP-Balanced Static Parameter PartitioningPipeline Payload ExtensionsPipeline ZeRO-2 Gradient Sharding with CPU OffloadingRedundant-Expert Capacity ReservationSConv-Aware Tensor-Parallel ShardingSequence Parallelism for Activations