PublishersvLLM
vLLM
10 documents read from this publisher.
Qwen3.8-Flash-Next on vLLM — Serve command for H200, H100, GB200 NVL4, GB300 NVL4, MI355X, RTX Pro 6000 4x, DGX Station (GB300)
software documentation · vendor docs · 2026-09-17 · 15 techniques
Mixture of ExpertsMulti-Token PredictionYaRNGated DeltaNetGated ResidualN-gram embedding host-memory offload and prefetchN-gram embedding lookupQwen Sparse AttentionTensor parallelism (degree 4)Expert parallelismGated DeltaNet and Qwen Sparse Attention hybridLimit maximum concurrent sequences to 256PLE CPU offload for N-gram embeddingsTensor and Expert Parallelism with Degree EightTiered CPU and filesystem offloading
zai-org/GLM-5.3 — 743B / 39B active · MOE · 1024K ctx
software documentation · vendor docs · 2026-09-10 · 9 techniques
zai-org/GLM-5.3-Flash — 321B / 18B active · MOE · 1024K ctx
software documentation · vendor docs · 2026-09-05 · 16 techniques
Mixture of ExpertsConfigurable reasoning effortMulti-Token PredictionFP8 KV-cache quantizationKimi Delta AttentionPrefill-decode disaggregationToken-level expert routingAlways-on thinking modeChat-template image/video placeholder expansionIdentical cache-layout pinning across prefill and decode poolsMXFP8 block-scale regrouping at loadNoPE sparse multi-head latent attentionReasoning-effort resolution and system-prompt injectionReusing the image token for video framesROCm AITER sparse MLA attention backendRound-robin routing for prefill-decode disaggregation
tencent/Hy4-preview — 770B / 49B active · MOE · 1024K ctx
software documentation · vendor docs · 2026-08-28 · 14 techniques
Mixture of ExpertsConfigurable reasoning effortMulti-Token PredictionFP8Top-8 expert routingAutomatic tool choiceGated DeepSeek Sparse AttentionIdentity Hyper-ConnectionsIndexCacheMoE with routed and shared expertsTensor parallelism (degree 8)FLASHMLA_SPARSEHy v4 tool-call parserReasoning parser hy_v4
thinkingmachines/Inkling — 975B / 41B active · MOE · 1024K ctx
software documentation · vendor docs · 2026-07-14 · 8 techniques
tencent/Hy3 — 295B / 21B active · MOE · 256K ctx
software documentation · vendor docs · 2026-07-06 · 15 techniques
Mixture of ExpertsMulti-Token PredictionSpeculative decodingFP8 KV-cache quantizationGrouped-query attentionShared ExpertsAITER fused-MoE backendAITER Linear backendAITER MHA backendAITER RMSNorm backendHPC-Ops attention backendHPC-Ops fused MoE backendRouted expertsTensorRT-LLM all-reduce backendTriton MoE kernel
Qwen/Qwen3.6-27B — 27B · DENSE · 256K ctx
software documentation · vendor docs · 2026-07-02 · 6 techniques
Qwen/Qwen3.6-35B-A3B — 35B / 3B active · MOE · 256K ctx
software documentation · vendor docs · 2026-06-03 · 18 techniques
Multi-Token PredictionYaRNFP8 KV-cache quantizationFP8Chunked prefillTensor ParallelismDisabling reasoning via chat-template configurationNVFP4 quantizationPrefix cachingAutomatic tool choiceAsynchronous schedulingCUDA Graph capture size reductionGated DeltaNet MoEfastsafetensorsMamba prefix caching in align modeMarlin MoE backendQwen3 reasoning parserQwen3 XML tool-call parser
XiaomiMiMo/MiMo-V2.5 — 311B / 15B active · MOE · 1024K ctx
software documentation · vendor docs · 2026-04-27 · 8 techniques
Qwen/Qwen3.5-397B-A17B — 397B / 17B active · MOE · 256K ctx
software documentation · vendor docs · 2026-04-16 · 12 techniques
YaRNFP8Expert ParallelismDisabling reasoning via chat-template configurationNVFP4 quantizationPrefix cachingLanguage-model-only serving modeData-parallel vision encodingMTP-1 speculative decodingreducing CUDA graph capture size to fit Mamba cacheshared-memory caching of preprocessed multimodal inputsTool calling