Model techniques map
Techniquesinference & servingdecoding strategy

implementation detail · filed under inference & serving

Concurrency-aware draft-length tuning

Adjusts speculative draft length with concurrency, with the reported optimal draft length decreasing as concurrency rises.

source
1
model
1
labs adopt it
0
strongest
evaluated

How sources treat it

One count per evidence span, weakest treatment to strongest.

evaluated 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

MTP is best suited for medium to high concurrency, with the optimal draft length decreasing as concurrency increases.

evaluatedinference servingin Nemotron 3.5 LightningNVIDIA

Filed alongside

Other methods under inference & serving :: decoding strategy.