Model techniques map
Techniquesinference & servingKV cache management

implementation detail · filed under inference & serving

Unified paged cache layout for KDA states and MLA KV

Places KDA states and MLA KV in a shared paged block pool with common allocation, reference-counting, and eviction machinery.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

We therefore pack KDA states into the same paged block pool as MLA KV , unifying pages to the same byte size so that both page types share one implementation of allocation, reference counting, and eviction.

coreinference servingin Kimi K3Moonshot AI

Filed alongside

Other methods under inference & serving :: KV cache management.