specific method · filed under inference & serving
RadixCache
A radix-based cache used for prefix sharing.
Also called radix cache.
- sources
- 2
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1not used 1
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
Context Management: Features like RadixCache (prefix sharing) and Prefix Cache
usedinference servingin GLM-5Z.ai
Unlike the commonly-used radix cache in current inference engines, our request-level prefix cache avoids re-prefilling or inter-request output cache sharing
not usedinference servingin MiMo-V2-FlashXiaomi
Filed alongside
Other methods under inference & serving :: KV cache management.
Chunked prefillPrefix cachingPrompt cachingSWA Bounded ReplayLanguage-model-only serving modePeriodic cache checkpointingCompressed KV cachingKDA-aware prefix-cache managementZero SWA caching8-bit Mamba cache quantizationAsynchronous cache offloading and restorationAutomatic cacheBlock-based KV cacheCross-layer KV-cache reuseCustomized heterogeneous KV-cache layoutDisable prefix caching for benchmarkingDynamic allocation of fixed-size state-cache poolsEncoder SWA bounded replayExact SWA KV reconstruction via full multi-layer replayFine-grained prefix hashingHierarchical KV cachingInference-side KV-cache reset on weight synchronizationKDA with prefill cacheKey-value reuse in global attention layers