specific method · filed under inference & serving
Fine-grained prefix hashing
Hashes prefixes at finer token blocks within coarser physical cache pages.
Also called Prefix hashing runs on fine hash blocks.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Prefix hashing runs on fine hash blocks (e.g., 512 tokens) inside MLA pages, while the physical block remains the coarse allocation unit.
coreinference servingin Kimi K3Moonshot AI
Filed alongside
Other methods under inference & serving :: KV cache management.
Chunked prefillPrefix cachingPrompt cachingSWA Bounded ReplayLanguage-model-only serving modePeriodic cache checkpointingCompressed KV cachingKDA-aware prefix-cache managementRadixCacheZero SWA caching8-bit Mamba cache quantizationAsynchronous cache offloading and restorationAutomatic cacheBlock-based KV cacheCross-layer KV-cache reuseCustomized heterogeneous KV-cache layoutDisable prefix caching for benchmarkingDynamic allocation of fixed-size state-cache poolsEncoder SWA bounded replayExact SWA KV reconstruction via full multi-layer replayHierarchical KV cachingInference-side KV-cache reset on weight synchronizationKDA with prefill cacheKey-value reuse in global attention layers