Model techniques map
Techniquesinference & servingKV cache management

specific method · filed under inference & serving

Fine-grained prefix hashing

Hashes prefixes at finer token blocks within coarser physical cache pages.

Also called Prefix hashing runs on fine hash blocks.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

Prefix hashing runs on fine hash blocks (e.g., 512 tokens) inside MLA pages, while the physical block remains the coarse allocation unit.

coreinference servingin Kimi K3Moonshot AI

Filed alongside

Other methods under inference & serving :: KV cache management.