specific method · filed under optimization
Slice-Granularity Elasticity
Maintaining continuous training with fewer TPU chip slices after a localized failure.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
we leverage Slice-Granularity Elasticity [Gemini Team, 2025], which allows continuous training with fewer “slices” of TPU chips when there is a localized failure.
usedsoftware implementationin Gemma 4Google DeepMind
Filed alongside
Other methods under optimization :: training runtime.
Offline checkpoint mergingAsynchronous checkpointingAsymmetric local-read and single-writer cache pathsAutotune configuration generationBatch-level embedding prefetchCaching the distributed checkpoint save planComposable activation storage policiesCPU-resident optimizer statesCross-rank remote activation offloadingDouble-buffered chunked streaming of reference-model weightsElement-wise activation recomputationGPU-to-GPU weight synchronization over GPUDirect RDMAHeartbeat-driven fault toleranceIn-flight recovery systemIn-flight weight updatesJob-level eviction and reclaimMinimal recomputation-graph extraction by backward traversalOverlapping NCCL transfers with device-to-host copiesPersistent checkpoint worker processesPersistent shared-storage cache for compiled artifactsPolicy-model gradient-buffer reuse for reference-model weightsPost-iteration NVMe offloading of training statesPre-admission hardware stress testingSame-node sticky pod respawn