ambiguous · filed under inference & serving
Prompt caching
Reuses cached prompt computation, though the evidence does not specify the underlying cache mechanism.
- sources
- 4
- models
- 4
- labs adopt it
- 3
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 4
Documented in
Evidence
4 spans quoted from the sources, strongest treatment first.
The chart below shows the average price customers are actually paying after prompt caching.
usedinference servingOpenRouter
The chart below shows the average price customers are actually paying after prompt caching.
usedinference servingin OpenRouterOpenRouter
The chart below shows the average price customers are actually paying after prompt caching.
usedinference serving
average price customers are actually paying after prompt caching
usedinference servingin OpenRouterOpenRouter
Filed alongside
Other methods under inference & serving :: KV cache management.
Chunked prefillPrefix cachingSWA Bounded ReplayLanguage-model-only serving modePeriodic cache checkpointingCompressed KV cachingKDA-aware prefix-cache managementRadixCacheZero SWA caching8-bit Mamba cache quantizationAsynchronous cache offloading and restorationAutomatic cacheBlock-based KV cacheCross-layer KV-cache reuseCustomized heterogeneous KV-cache layoutDisable prefix caching for benchmarkingDynamic allocation of fixed-size state-cache poolsEncoder SWA bounded replayExact SWA KV reconstruction via full multi-layer replayFine-grained prefix hashingHierarchical KV cachingInference-side KV-cache reset on weight synchronizationKDA with prefill cacheKey-value reuse in global attention layers