Model techniques map
Techniquesinference & servinginference scheduling

specific method · filed under inference & serving

Prefix-cache-aware session affinity scheduling

A scheduling method that routes sessions to clusters holding their prefix cache while bounding cluster-failure costs.

Also called cache-aware affinity scheduling.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

cache-aware affinity scheduling routes each session to the cluster holding its prefix cache while bounding the cost of cluster failures

coreinference servingin Kimi K3Moonshot AI

Filed alongside

Other methods under inference & serving :: inference scheduling.