ambiguous · filed under inference & serving
Compressed KV caching
Reduces KV-cache storage requirements, but the evidence does not specify a single compression mechanism.
Also called KV cache compression, Smaller KV cache.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
Compared with the previous generation, V4.1-Flash's KV cache needs just: 1/4 the HBM, 1/8 the SSD storage
coreinference servingin DeepSeek-V4.1-FlashDeepSeek
Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation
coreinference servingin DeepSeek-V4.1-FlashDeepSeek
Filed alongside
Other methods under inference & serving :: KV cache management.
Chunked prefillPrefix cachingPrompt cachingSWA Bounded ReplayLanguage-model-only serving modePeriodic cache checkpointingKDA-aware prefix-cache managementRadixCacheZero SWA caching8-bit Mamba cache quantizationAsynchronous cache offloading and restorationAutomatic cacheBlock-based KV cacheCross-layer KV-cache reuseCustomized heterogeneous KV-cache layoutDisable prefix caching for benchmarkingDynamic allocation of fixed-size state-cache poolsEncoder SWA bounded replayExact SWA KV reconstruction via full multi-layer replayFine-grained prefix hashingHierarchical KV cachingInference-side KV-cache reset on weight synchronizationKDA with prefill cacheKey-value reuse in global attention layers