Model techniques map
Taxonomyinference & servingKV cache management

taxonomy node · level 2

KV cache management

42 methods filed at this node or below it, from the sources of 17 models.

inference & serving :: KV cache management

Matching aids for the classifier: paged KV cache; prefix caching; radix cache; chunked prefill; periodic cache checkpointing; block-based key-value cache.

In this branch 42

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 42

Compressed KV caching core · 2 sources · 2 quotes
SWA Bounded Replay core · 2 sources · 4 quotes
Block-based KV cache core · 1 source · 1 quote
Cross-layer KV-cache reuse core · 1 source · 1 quote
Customized heterogeneous KV-cache layout core · 1 source · 1 quote
Fine-grained prefix hashing core · 1 source · 1 quote
KDA-aware prefix-cache management core · 1 source · 2 quotes
Persistent per-dialogue-context KV caching core · 1 source · 1 quote
Shared-free-list cache allocation core · 1 source · 1 quote
Write-back external KV-cache policy core · 1 source · 1 quote
Chunked prefill default · 6 sources · 6 quotes
Automatic cache default · 1 source · 1 quote
Prefix caching used · 5 sources · 5 quotes
Prompt caching used · 4 sources · 4 quotes
Language-model-only serving mode used · 3 sources · 3 quotes
RadixCache used · 2 sources · 2 quotes
Disable prefix caching for benchmarking used · 1 source · 1 quote
Encoder SWA bounded replay used · 1 source · 1 quote
Hierarchical KV caching used · 1 source · 1 quote
KDA with prefill cache used · 1 source · 1 quote
Key-value reuse in global attention layers used · 1 source · 1 quote
KV-cache sharing used · 1 source · 1 quote
Least-recently-used eviction used · 1 source · 1 quote
On-disk KV-cache storage used · 1 source · 1 quote
Overlapped host-memory prefetching used · 1 source · 1 quote
Persistent KV-cache management used · 1 source · 1 quote
Pre-scheduled tile chunking used · 1 source · 1 quote
Request-level prefix cache used · 1 source · 1 quote
Stateless in-memory prefix caching used · 1 source · 1 quote
Periodic cache checkpointing optional · 3 sources · 3 quotes
Zero SWA caching optional · 2 sources · 2 quotes
8-bit Mamba cache quantization evaluated · 1 source · 1 quote
Mamba prefix caching in align mode evaluated · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modeltechniques
DeepSeek-V4.1-Flash coreCompressed KV caching coreCross-layer KV-cache reuse coreSWA Bounded Replay coreEncoder SWA bounded replay usedLeast-recently-used eviction usedPersistent KV-cache management usedExact SWA KV reconstruction via full multi-layer replay not usedZero SWA caching not used—
NVIDIA-Nemotron-3-Ultra-550B-A55B defaultChunked prefill defaultPrefix caching used8-bit Mamba cache quantization evaluatedPeriodic cache checkpointing evaluated—
MiMo-V2.6-Flash corePersistent per-dialogue-context KV caching coreAsynchronous cache offloading and restoration usedHierarchical KV caching used—
DeepSeek-V4-Flash coreCustomized heterogeneous KV-cache layout coreDynamic allocation of fixed-size state-cache pools usedKV-cache and sparse-attention-kernel co-design usedOn-disk KV-cache storage usedPeriodic cache checkpointing optionalZero SWA caching optional—
MiMo-V2.5 usedChunked prefill usedRequest-level prefix cache usedRadixCache not used—
GLM-5.3 usedDisable prefix caching for benchmarking used—
GLM-5.2 usedPrefix caching usedRadixCache used—
MiniMax-M3 coreBlock-based KV cache coreAutomatic cache defaultPre-scheduled tile chunking used—
DeepSeek-V4-Pro coreCustomized heterogeneous KV-cache layout coreDynamic allocation of fixed-size state-cache pools usedKV-cache and sparse-attention-kernel co-design usedOn-disk KV-cache storage usedPeriodic cache checkpointing optionalZero SWA caching optional—
Gemma 4 31B usedKey-value reuse in global attention layers usedKV-cache sharing usedPrompt caching usedStateless in-memory prefix caching used—
Inkling usedPrompt caching used—
Kimi K3 coreFine-grained prefix hashing coreKDA-aware prefix-cache management coreProjected-input caching for speculative KDA rollback coreShared-free-list cache allocation coreSparse hash-aligned KDA recurrent-state checkpoints coreUnified paged cache layout for KDA states and MLA KV coreWrite-back external KV-cache policy coreKDA with prefill cache usedKV-cache-aware placement of one-shot option messages used—
Laguna-S-2.1 usedInference-side KV-cache reset on weight synchronization used—
MiMo-V2.6-Pro corePersistent per-dialogue-context KV caching coreAsynchronous cache offloading and restoration usedHierarchical KV caching used—
Qwen3.5-397B-A17B usedChunked prefill usedLanguage-model-only serving mode usedPrompt caching usedPrefix caching optional—
Qwen3.6-35B-A3B usedChunked prefill usedPrefix caching usedPrompt caching usedLanguage-model-only serving mode optionalMamba prefix caching in align mode evaluated—
Qwen3.8-Flash-Next usedOverlapped host-memory prefetching used—