specific method · filed under optimization
In-flight recovery system
An automated system that detects training failures, replaces affected workers, and resumes from the latest checkpoint.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
an in-flight recovery system detects hangs, NCCL failures, and node losses, replaces affected workers, and resumes from the latest checkpoint without human involvement.
coresoftware implementationin Model FactoryPoolside
Filed alongside
Other methods under optimization :: training runtime.
Offline checkpoint mergingAsynchronous checkpointingAsymmetric local-read and single-writer cache pathsAutotune configuration generationBatch-level embedding prefetchCaching the distributed checkpoint save planComposable activation storage policiesCPU-resident optimizer statesCross-rank remote activation offloadingDouble-buffered chunked streaming of reference-model weightsElement-wise activation recomputationGPU-to-GPU weight synchronization over GPUDirect RDMAHeartbeat-driven fault toleranceIn-flight weight updatesJob-level eviction and reclaimMinimal recomputation-graph extraction by backward traversalOverlapping NCCL transfers with device-to-host copiesPersistent checkpoint worker processesPersistent shared-storage cache for compiled artifactsPolicy-model gradient-buffer reuse for reference-model weightsPost-iteration NVMe offloading of training statesPre-admission hardware stress testingSame-node sticky pod respawnSeeding node-local storage from a warm shared cache