Model techniques map
Techniquesinference & servingKV cache management

ambiguous · filed under inference & serving

Prompt caching

Reuses cached prompt computation, though the evidence does not specify the underlying cache mechanism.

sources
4
models
4
labs adopt it
3
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 4

Documented in

Evidence

4 spans quoted from the sources, strongest treatment first.

The chart below shows the average price customers are actually paying after prompt caching.

usedinference servingOpenRouter

The chart below shows the average price customers are actually paying after prompt caching.

usedinference servingin OpenRouterOpenRouter

The chart below shows the average price customers are actually paying after prompt caching.

usedinference serving

average price customers are actually paying after prompt caching

usedinference servingin OpenRouterOpenRouter

Filed alongside

Other methods under inference & serving :: KV cache management.