specific method · filed under software implementation
vLLM
An inference and serving engine used to deploy models, with recipes and operational support described in the evidence.
Also called Serving MiMo-V2.6-Pro-RL with vLLM, vLLM model serving, serves the model with vLLM.
- sources
- 7
- models
- 5
- labs adopt it
- 5
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Further reading
Picked by hand, not extracted: where to read more, not evidence for anything on this page.
- vLLM Documentation docs docs.vllm.ai
Evidence
7 spans quoted from the sources, strongest treatment first.
disaggregation for hybrid Mamba-Attention models now works out-of-the-box in vLLM (Kwon et al., 2023)
Acceleration Engine: vLLM
disaggregation for hybrid Mamba-Attention models now works out-of-the-box in vLLM
Follow the vLLM MiMo-V2.5 recipe.
We recommend the following inference frameworks to serve the model: vLLM
the command below serves the model with vLLM on a single 4×GB300 node.
vLLM (v0.23.0+)
Filed alongside
Other methods under software implementation :: inference engine.