implementation detail · filed under inference & serving
KV-cache sharing between drafter and target
Lets the draft model reuse target-model activations and its KV cache instead of recalculating context.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
The draft models seamlessly utilize the target model's activations and share its KV cache, meaning they don't have to waste time recalculating context the larger model has already figured out.
usedinference servingin Gemma 4 MTP draftersGoogle
Filed alongside
Other methods under inference & serving :: decoding strategy.
Speculative decodingMulti-Token PredictionDSparkEAGLEDFlashPresence PenaltyBest-of-N scaffoldingNEXTN speculative decodingMulti-layer EAGLESpeculative samplingStandardized sampling configurationMulti-stage candidate filteringRecursive shared MTP-head draftingTask-specific sampling parametersChat Prefix CompletionConcurrency-aware draft-length tuningDistribution-matched draft-model fine-tuningEAGLE-3-style draft-model fine-tuningFused recurrent replay kernelGrammar-constrained decodingLongest-trace selectionMTP-1 speculative decodingSame-checkpoint target and draft weightsThroughput-aware dynamic verification-length scheduling