specific method · filed under optimization
Policy-model gradient-buffer reuse for reference-model weights
Backing reference-model parameter tensors with the policy model's FP32 gradient-buffer storage and materializing them when needed.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
We keep these weights in CPU memory and materialize them only when needed, backing their parameter tensors with the policy model’s FP32 gradient-buffer storage.
usedpost trainingin Kimi K3Moonshot AI
Filed alongside
Other methods under optimization :: training runtime.
Offline checkpoint mergingAsynchronous checkpointingAsymmetric local-read and single-writer cache pathsAutotune configuration generationBatch-level embedding prefetchCaching the distributed checkpoint save planComposable activation storage policiesCPU-resident optimizer statesCross-rank remote activation offloadingDouble-buffered chunked streaming of reference-model weightsElement-wise activation recomputationGPU-to-GPU weight synchronization over GPUDirect RDMAHeartbeat-driven fault toleranceIn-flight recovery systemIn-flight weight updatesJob-level eviction and reclaimMinimal recomputation-graph extraction by backward traversalOverlapping NCCL transfers with device-to-host copiesPersistent checkpoint worker processesPersistent shared-storage cache for compiled artifactsPost-iteration NVMe offloading of training statesPre-admission hardware stress testingSame-node sticky pod respawnSeeding node-local storage from a warm shared cache