Model techniques map
Taxonomyoptimizationtraining runtime

taxonomy node · level 2

training runtime

29 methods filed at this node or below it, from the sources of 11 models.

optimization :: training runtime

Matching aids for the classifier: asynchronous checkpointing; in-flight weight updates; offline checkpoint merging; heartbeat-driven fault tolerance.

In this branch 29

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 29

Composable activation storage policies core · 1 source · 1 quote
In-flight recovery system core · 1 source · 1 quote
Offline checkpoint merging used · 3 sources · 3 quotes
Asynchronous checkpointing used · 2 sources · 2 quotes
Autotune configuration generation used · 1 source · 1 quote
Batch-level embedding prefetch used · 1 source · 1 quote
CPU-resident optimizer states used · 1 source · 1 quote
Cross-rank remote activation offloading used · 1 source · 1 quote
Element-wise activation recomputation used · 1 source · 1 quote
Heartbeat-driven fault tolerance used · 1 source · 1 quote
In-flight weight updates used · 1 source · 1 quote
Job-level eviction and reclaim used · 1 source · 1 quote
Persistent checkpoint worker processes used · 1 source · 1 quote
Pre-admission hardware stress testing used · 1 source · 1 quote
Same-node sticky pod respawn used · 1 source · 1 quote
Slice-Granularity Elasticity used · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modeltechniques
DeepSeek-V4.1-Flash usedBatch-level embedding prefetch usedOffline checkpoint merging used—
NVIDIA-Nemotron-3-Ultra-550B-A55B usedAsymmetric local-read and single-writer cache paths usedAsynchronous checkpointing usedCaching the distributed checkpoint save plan usedIn-flight weight updates usedOffline checkpoint merging usedOverlapping NCCL transfers with device-to-host copies usedPersistent checkpoint worker processes usedPersistent shared-storage cache for compiled artifacts usedSeeding node-local storage from a warm shared cache used—
MiMo-V2.6-Flash usedCPU-resident optimizer states used—
DeepSeek-V4-Flash coreTensor-level automatic-differentiation activation checkpointing coreMinimal recomputation-graph extraction by backward traversal usedSelective tensor recomputation for activation-memory reduction used—
GLM-5.2 usedHeartbeat-driven fault tolerance used—
MiniMax-M3 usedAutotune configuration generation used—
DeepSeek-V4-Pro coreTensor-level automatic-differentiation activation checkpointing coreMinimal recomputation-graph extraction by backward traversal usedSelective tensor recomputation for activation-memory reduction used—
Gemma 4 31B usedSlice-Granularity Elasticity used—
Kimi K3 coreComposable activation storage policies coreUnified pluggable activation storage manager coreCross-rank remote activation offloading usedDouble-buffered chunked streaming of reference-model weights usedElement-wise activation recomputation usedPolicy-model gradient-buffer reuse for reference-model weights usedPost-iteration NVMe offloading of training states used—
Laguna-S-2.1 coreIn-flight recovery system coreGPU-to-GPU weight synchronization over GPUDirect RDMA usedJob-level eviction and reclaim usedPre-admission hardware stress testing usedSame-node sticky pod respawn used—
MiMo-V2.6-Pro usedCPU-resident optimizer states used—