ambiguous · filed under inference & serving
Periodic cache checkpointing
Saves cache or SWA KV checkpoints periodically; the evidence does not fully specify the checkpoint interval or all checkpointed state.
Also called cache checkpointing, Periodic checkpointing of SWA KV entries, Periodic Checkpointing.
- sources
- 3
- models
- 3
- lab adopt it
- 1
- strongest
- optional
How sources treat it
One count per evidence span, weakest treatment to strongest.
evaluated 2optional 1
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
This strategy checkpoints SWA KV entries of the last tokens within every tokens, where is a tunable parameter.
optionalinference servingin DeepSeek-V4DeepSeek
We ran initial experiments on Nemotron 3 Super NVFP4, using emulated quantization
evaluatedpost trainingin Nemotron 3 SuperNVIDIA
we explore the idea of periodic cache checkpointing
evaluatedinference servingin Nemotron 3 SuperNVIDIA
Filed alongside
Other methods under inference & serving :: KV cache management.
Chunked prefillPrefix cachingPrompt cachingSWA Bounded ReplayLanguage-model-only serving modeCompressed KV cachingKDA-aware prefix-cache managementRadixCacheZero SWA caching8-bit Mamba cache quantizationAsynchronous cache offloading and restorationAutomatic cacheBlock-based KV cacheCross-layer KV-cache reuseCustomized heterogeneous KV-cache layoutDisable prefix caching for benchmarkingDynamic allocation of fixed-size state-cache poolsEncoder SWA bounded replayExact SWA KV reconstruction via full multi-layer replayFine-grained prefix hashingHierarchical KV cachingInference-side KV-cache reset on weight synchronizationKDA with prefill cacheKey-value reuse in global attention layers