specific method · filed under optimization
Batch-level embedding prefetch
Prefetching embeddings for the entire local batch before pipeline stages process their microbatches.
Also called Embedding prefetch.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Embedding prefetch is therefore initiated for the entire local batch before each pipeline stage begins processing microbatches
usedunclearin DeepSeek-V4.1-FlashDeepSeek
Filed alongside
Other methods under optimization :: training runtime.
Offline checkpoint mergingAsynchronous checkpointingAsymmetric local-read and single-writer cache pathsAutotune configuration generationCaching the distributed checkpoint save planComposable activation storage policiesCPU-resident optimizer statesCross-rank remote activation offloadingDouble-buffered chunked streaming of reference-model weightsElement-wise activation recomputationGPU-to-GPU weight synchronization over GPUDirect RDMAHeartbeat-driven fault toleranceIn-flight recovery systemIn-flight weight updatesJob-level eviction and reclaimMinimal recomputation-graph extraction by backward traversalOverlapping NCCL transfers with device-to-host copiesPersistent checkpoint worker processesPersistent shared-storage cache for compiled artifactsPolicy-model gradient-buffer reuse for reference-model weightsPost-iteration NVMe offloading of training statesPre-admission hardware stress testingSame-node sticky pod respawnSeeding node-local storage from a warm shared cache