specific method · filed under optimization
Offline checkpoint merging
Merging checkpoints offline for evaluation analysis or to combine improvements across runs with different scaffolds or configurations.
- sources
- 3
- models
- 2
- labs adopt it
- 2
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 3
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
offline checkpoint merging for evaluation analysis
usedoptimizationin Nemotron 3 UltraNVIDIA
we use model merging to reinitialize successive RL runs. Specifically, we merge checkpoints from runs across different scaffolds or configurations, combining improvements acquired along different optimization paths.
usedpost trainingin DeepSeek-V4.1DeepSeek
we used offline checkpoint merging for evaluation analysis
usedoptimizationin Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under optimization :: training runtime.
Asynchronous checkpointingAsymmetric local-read and single-writer cache pathsAutotune configuration generationBatch-level embedding prefetchCaching the distributed checkpoint save planComposable activation storage policiesCPU-resident optimizer statesCross-rank remote activation offloadingDouble-buffered chunked streaming of reference-model weightsElement-wise activation recomputationGPU-to-GPU weight synchronization over GPUDirect RDMAHeartbeat-driven fault toleranceIn-flight recovery systemIn-flight weight updatesJob-level eviction and reclaimMinimal recomputation-graph extraction by backward traversalOverlapping NCCL transfers with device-to-host copiesPersistent checkpoint worker processesPersistent shared-storage cache for compiled artifactsPolicy-model gradient-buffer reuse for reference-model weightsPost-iteration NVMe offloading of training statesPre-admission hardware stress testingSame-node sticky pod respawnSeeding node-local storage from a warm shared cache