taxonomy node · level 2
inference engine
15 methods filed at this node or below it, from the sources of 13 models.
software implementation :: inference engine
Matching aids for the classifier: vLLM; SGLang; TensorRT-LLM; llama.cpp; KTransformers; xLLM; vLLM-Ascend.
In this branch 15
Everything filed at this node or below it, with one collapsible heading per child node.
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| Hy4-preview used | vLLM or SGLang serving used— |
| NVIDIA-Nemotron-3-Ultra-550B-A55B used | Environment variable propagation to subprocesses usedvLLM usedSGLang optionalTensorRT-LLM optional— |
| MiMo-V2.6-Flash used | vLLM tensor parallelism used— |
| GLM-5.2 optional | KTransformers optionalSGLang optionalUnsloth optionalvLLM optionalvLLM-Ascend optionalxLLM optional— |
| MiniMax-M3 optional | KTransformers optionalSGLang optionalUnsloth optionalvLLM optional— |
| DeepSeek-V4-Flash-Vision-Exp optional | vLLM optional— |
| Gemma 4 31B used | litert-lm serve used— |
| Laguna-S-2.1 used | Atlas inference library used— |
| MiMo-V2.6-Pro used | vLLM tensor parallelism usedSGLang optionalvLLM optional— |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B used | TensorRT-LLM used— |
| Qwen3.5-397B-A17B optional | vLLM language-model-only mode optional— |
| Qwen3.6-35B-A3B optional | Dedicated serving engines optional— |
| Step-3.7-Flash used | llama.cpp used— |