specific method · filed under inference & serving
Prefix-cache-aware session affinity scheduling
A scheduling method that routes sessions to clusters holding their prefix cache while bounding cluster-failure costs.
Also called cache-aware affinity scheduling.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
cache-aware affinity scheduling routes each session to the cluster holding its prefix cache while bounding the cost of cluster failures
coreinference servingin Kimi K3Moonshot AI
Filed alongside
Other methods under inference & serving :: inference scheduling.
Asynchronous schedulingNUMA binding of workers to GPU-local CPU socketsComposite-key sorting for topology-aware GPU rank assignmentContinuous batchingCross-group pinning of cache-hit blocksDual-cluster consistent-hash failover for cache affinityEnvoy-based proxy with custom orchestratorExacto routingGreedy feasible-rank placement by remaining capacityHost-side scheduling optimizationIncreasing batch size for inferenceLatency-sensitive execution class with priority isolationMax sequences tuningPer-node hard admission constraintRegistering NVLink domains as Ray custom resourcesRequest-class resource-budget admission controlRuntime-signal-based rollout concurrency auto-throttlingSub-NUMA partitioning with per-VM NUMA bindingWorkload-aware routed-expert GEMM scheduling