specific method · filed under optimization
Asynchronous checkpointing
Checkpointing that copies model parameters to CPU and persists them in the background while training continues.
Also called Asynchronous background checkpointing.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
Enabling asynchronous checkpointing from the Nvidia Resiliency Extension (NVRx), where model parameters are copied to CPU and persisted in the background while training continues
usedsoftware implementationin Nemotron 3 UltraNVIDIA
Enabling asynchronous checkpointing from the Nvidia Resiliency Extension (NVRx)
usedsoftware implementationin Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under optimization :: training runtime.
Offline checkpoint mergingAsymmetric local-read and single-writer cache pathsAutotune configuration generationBatch-level embedding prefetchCaching the distributed checkpoint save planComposable activation storage policiesCPU-resident optimizer statesCross-rank remote activation offloadingDouble-buffered chunked streaming of reference-model weightsElement-wise activation recomputationGPU-to-GPU weight synchronization over GPUDirect RDMAHeartbeat-driven fault toleranceIn-flight recovery systemIn-flight weight updatesJob-level eviction and reclaimMinimal recomputation-graph extraction by backward traversalOverlapping NCCL transfers with device-to-host copiesPersistent checkpoint worker processesPersistent shared-storage cache for compiled artifactsPolicy-model gradient-buffer reuse for reference-model weightsPost-iteration NVMe offloading of training statesPre-admission hardware stress testingSame-node sticky pod respawnSeeding node-local storage from a warm shared cache