implementation detail · filed under inference & serving
Concurrency-aware draft-length tuning
Adjusts speculative draft length with concurrency, with the reported optimal draft length decreasing as concurrency rises.
- source
- 1
- model
- 1
- labs adopt it
- 0
- strongest
- evaluated
How sources treat it
One count per evidence span, weakest treatment to strongest.
evaluated 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
MTP is best suited for medium to high concurrency, with the optimal draft length decreasing as concurrency increases.
evaluatedinference servingin Nemotron 3.5 LightningNVIDIA
Filed alongside
Other methods under inference & serving :: decoding strategy.
Speculative decodingMulti-Token PredictionDSparkEAGLEDFlashPresence PenaltyBest-of-N scaffoldingNEXTN speculative decodingMulti-layer EAGLESpeculative samplingStandardized sampling configurationMulti-stage candidate filteringRecursive shared MTP-head draftingTask-specific sampling parametersChat Prefix CompletionDistribution-matched draft-model fine-tuningEAGLE-3-style draft-model fine-tuningFused recurrent replay kernelGrammar-constrained decodingKV-cache sharing between drafter and targetLongest-trace selectionMTP-1 speculative decodingSame-checkpoint target and draft weightsThroughput-aware dynamic verification-length scheduling