general family · filed under inference & serving
Task-specific sampling parameters
Uses different recommended sampling-parameter sets according to task or generation mode.
Also called sampling parameters.
- sources
- 2
- models
- 2
- lab adopt it
- 1
- strongest
- optional
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
We recommend using the following set of sampling parameters for generation
optionalinference servingin Qwen3.5-122B-A10BQwen
We suggest using the following sets of sampling parameters depending on the mode and task type
optionalinference servingin Qwen3.6-27BQwen
Filed alongside
Other methods under inference & serving :: decoding strategy.
Speculative decodingMulti-Token PredictionDSparkEAGLEDFlashPresence PenaltyBest-of-N scaffoldingNEXTN speculative decodingMulti-layer EAGLESpeculative samplingStandardized sampling configurationMulti-stage candidate filteringRecursive shared MTP-head draftingChat Prefix CompletionConcurrency-aware draft-length tuningDistribution-matched draft-model fine-tuningEAGLE-3-style draft-model fine-tuningFused recurrent replay kernelGrammar-constrained decodingKV-cache sharing between drafter and targetLongest-trace selectionMTP-1 speculative decodingSame-checkpoint target and draft weightsThroughput-aware dynamic verification-length scheduling