specific method · filed under inference & serving
NUMA binding of workers to GPU-local CPU sockets
A placement method that binds worker processes to the CPU socket local to their assigned GPUs.
Also called NUMA Binding for CPU-GPU affinity, NUMA Binding.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
The fix was to explicitly bind policy and vLLM worker processes to the CPU socket local to their assigned GPUs.
usedsoftware implementationin NeMo-RLNVIDIA
The fix was to explicitly bind policy and vLLM worker processes to the CPU socket local to their assigned GPUs.
usedsoftware implementationin Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under inference & serving :: inference scheduling.
Asynchronous schedulingComposite-key sorting for topology-aware GPU rank assignmentContinuous batchingCross-group pinning of cache-hit blocksDual-cluster consistent-hash failover for cache affinityEnvoy-based proxy with custom orchestratorExacto routingGreedy feasible-rank placement by remaining capacityHost-side scheduling optimizationIncreasing batch size for inferenceLatency-sensitive execution class with priority isolationMax sequences tuningPer-node hard admission constraintPrefix-cache-aware session affinity schedulingRegistering NVLink domains as Ray custom resourcesRequest-class resource-budget admission controlRuntime-signal-based rollout concurrency auto-throttlingSub-NUMA partitioning with per-VM NUMA bindingWorkload-aware routed-expert GEMM scheduling