implementation detail · filed under inference & serving
CUDA Graph capture size reduction
Reducing the CUDA Graph capture-size limit using the max-cudagraph-capture-size setting.
Also called reduce --max-cudagraph-capture-size.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 1used 1
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
reduce --max-cudagraph-capture-size(default 512)
usedunclearin Qwen3.6-35B-A3BQwen
CUDA graph / Mamba cache size error: reduce --max-cudagraph-capture-size(default 512).
optionalinference servingin vLLMvLLM
Filed alongside
Other methods under inference & serving :: inference kernel.
Batch-invariant deterministic kernelsCUDA GraphFlashAttention 3Kernel fusionMoE-side chunkingSynchronization-free static-shape MoE executionTensorRT-LLM multi-head attention backendAvoiding split-KDeepGEMM-based batch-invariant matrix multiplicationDistributed shared memory for cross-SM attention data exchangeDual-kernel batch-invariant attention decodingDynamic load balancingExp-free TopK kernelExpert-optimized Triton kernelsFA4 sheared-bias attention kernelFlashAttentionFlashAttention 4FlashKDAFP8 GEMMFused AttnRes merge and RMSNorm kernelFused latent down-projection and MoE-router GEMMFused mHC kernelsFused QSA kernelFused RoPE-attention-RoPE-cast kernel