specific method · filed under software implementation
SGLang
An inference and serving framework with model-specific deployment guidance.
Also called Serving MiMo-V2.6-Pro-RL with SGLang.
- sources
- 4
- models
- 4
- labs adopt it
- 4
- strongest
- optional
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 4
Documented in
Further reading
Picked by hand, not extracted: where to read more, not evidence for anything on this page.
- SGLang: Efficient Execution of Structured Language Model Programs (Zheng et al., 2023) paper arxiv.orgintroduces SGLang and RadixAttention
- sgl-project/sglang code github.comofficial inference-serving framework repository
Evidence
4 spans quoted from the sources, strongest treatment first.
For best performance, follow the SGLang MiMo cookbook.
optionalinference servingin MiMo-V2.6-Pro-RLXiaomi
We recommend the following inference frameworks to serve the model: SGLang
optionalsoftware implementationin MiniMax-M3MiniMax
SGLang deployment snippet provided
optionalsoftware implementationin Nemotron 3 UltraNVIDIA
SGLang (v0.5.13.post1+)
optionalsoftware implementationin GLM-5.2Z.ai
Filed alongside
Other methods under software implementation :: inference engine.