Model techniques map
Techniquessoftware implementationinference engine

specific method · filed under software implementation

vLLM

An inference and serving engine used to deploy models, with recipes and operational support described in the evidence.

Also called Serving MiMo-V2.6-Pro-RL with vLLM, vLLM model serving, serves the model with vLLM.

sources
7
models
5
labs adopt it
5
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

optional 4used 3

Documented in

Further reading

Picked by hand, not extracted: where to read more, not evidence for anything on this page.

Evidence

7 spans quoted from the sources, strongest treatment first.

disaggregation for hybrid Mamba-Attention models now works out-of-the-box in vLLM (Kwon et al., 2023)

usedsoftware implementationin Nemotron 3 Ultra

Acceleration Engine: vLLM

usedsoftware implementationin Nemotron 3 UltraNVIDIA

disaggregation for hybrid Mamba-Attention models now works out-of-the-box in vLLM

usedsoftware implementationin Nemotron 3 UltraNVIDIA

Follow the vLLM MiMo-V2.5 recipe.

optionalinference servingin MiMo-V2.6-Pro-RLXiaomi

We recommend the following inference frameworks to serve the model: vLLM

optionalsoftware implementationin MiniMax-M3MiniMax

the command below serves the model with vLLM on a single 4×GB300 node.

optionalinference servingin DeepSeek-V4-Flash-Vision-ExpDeepSeek

vLLM (v0.23.0+)

optionalsoftware implementationin GLM-5.2Z.ai

Filed alongside

Other methods under software implementation :: inference engine.