specific method · filed under inference & serving
Multi-layer EAGLE
An EAGLE variant enabled as a multi-layer configuration; no further mechanism is specified in the evidence.
Also called enable-multi-layer-eagle.
- sources
- 3
- models
- 2
- labs adopt it
- 2
- strongest
- optional
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 3
Documented in
Further reading
Picked by hand, not extracted: where to read more, not evidence for anything on this page.
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test (Li et al., 2025) paper arxiv.orgmulti-layer feature fusion draft model
- Speculative Decoding - SGLang Documentation docs docs.sglang.aidocuments the multi-layer EAGLE worker option
Evidence
3 spans quoted from the sources, strongest treatment first.
--enable-multi-layer-eagle
optionalinference servingin MiMo-V2.5-ProXiaomi
--enable-multi-layer-eagle
optionalinference servingin Step 3.7 FlashStepFun
Filed alongside
Other methods under inference & serving :: decoding strategy.
Speculative decodingMulti-Token PredictionDSparkEAGLEDFlashPresence PenaltyBest-of-N scaffoldingNEXTN speculative decodingSpeculative samplingStandardized sampling configurationMulti-stage candidate filteringRecursive shared MTP-head draftingTask-specific sampling parametersChat Prefix CompletionConcurrency-aware draft-length tuningDistribution-matched draft-model fine-tuningEAGLE-3-style draft-model fine-tuningFused recurrent replay kernelGrammar-constrained decodingKV-cache sharing between drafter and targetLongest-trace selectionMTP-1 speculative decodingSame-checkpoint target and draft weightsThroughput-aware dynamic verification-length scheduling