Model techniques map
Techniquesinference & servinginference scheduling

implementation detail · filed under inference & serving

Cross-group pinning of cache-hit blocks

A cache-consistency detail that pins each hit block across all groups before allocation.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

every hit block is therefore pinned across all groups before anything is allocated.

coreinference servingin Kimi K3Moonshot AI

Filed alongside

Other methods under inference & serving :: inference scheduling.