specific method · filed under inference & serving
Avoiding split-K
A batch-invariance strategy that abandons split-K in most scenarios, potentially at a performance cost.
Also called we abandon split-k in most scenarios.
- source
- 1
- models
- 2
- labs adopt it
- 0
- strongest
- not used
How sources treat it
One count per evidence span, weakest treatment to strongest.
not used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Therefore, we abandon split-k in most scenarios, which, however, may cause performance degradation.
not usedunclearin DeepSeek-V4DeepSeek
Filed alongside
Other methods under inference & serving :: inference kernel.
Batch-invariant deterministic kernelsCUDA GraphCUDA Graph capture size reductionFlashAttention 3Kernel fusionMoE-side chunkingSynchronization-free static-shape MoE executionTensorRT-LLM multi-head attention backendDeepGEMM-based batch-invariant matrix multiplicationDistributed shared memory for cross-SM attention data exchangeDual-kernel batch-invariant attention decodingDynamic load balancingExp-free TopK kernelExpert-optimized Triton kernelsFA4 sheared-bias attention kernelFlashAttentionFlashAttention 4FlashKDAFP8 GEMMFused AttnRes merge and RMSNorm kernelFused latent down-projection and MoE-router GEMMFused mHC kernelsFused QSA kernelFused RoPE-attention-RoPE-cast kernel