Model techniques map
Taxonomyinference & serving

taxonomy area · pipeline stage 06

inference & serving

375 methods filed at this node or below it, from the sources of 28 models.

inference & serving

Matching aids for the classifier: how a finished model is run.

In this branch 375

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 1

communication optimization used · 1 source · 1 quote

decoding strategy 28

Speculative decoding core · 22 sources · 25 quotes
Multi-Token Prediction core · 20 sources · 20 quotes
DSpark core · 8 sources · 10 quotes
DFlash core · 6 sources · 8 quotes
Fused recurrent replay kernel core · 1 source · 1 quote
Recursive shared MTP-head drafting core · 1 source · 2 quotes
Standardized sampling configuration default · 3 sources · 3 quotes
Same-checkpoint target and draft weights default · 1 source · 1 quote
EAGLE used · 9 sources · 9 quotes
NEXTN speculative decoding used · 4 sources · 4 quotes
Presence Penalty used · 4 sources · 5 quotes
Best-of-N scaffolding used · 3 sources · 4 quotes
Multi-stage candidate filtering used · 2 sources · 2 quotes
EAGLE-3-style draft-model fine-tuning used · 1 source · 1 quote
Grammar-constrained decoding used · 1 source · 1 quote
Longest-trace selection used · 1 source · 1 quote
MTP-1 speculative decoding used · 1 source · 1 quote
Throughput-based draft-block sizing used · 1 source · 1 quote
Top-k over token clusters used · 1 source · 1 quote
Top-p and top-k sampling used · 1 source · 1 quote
Multi-layer EAGLE optional · 3 sources · 3 quotes
Speculative sampling optional · 3 sources · 3 quotes
Task-specific sampling parameters optional · 2 sources · 2 quotes
Chat Prefix Completion optional · 1 source · 1 quote
Concurrency-aware draft-length tuning evaluated · 1 source · 1 quote

reasoning control 42

Configurable reasoning effort core · 33 sources · 34 quotes
Chain-of-thought reasoning core · 8 sources · 8 quotes
Default thinking mode core · 8 sources · 8 quotes
Interleaved thinking between tool calls core · 3 sources · 4 quotes
Always-on thinking mode core · 2 sources · 2 quotes
Control-token-enabled thinking mode core · 2 sources · 2 quotes
Thinking mode selection default · 4 sources · 4 quotes
Maximum thinking effort default · 3 sources · 3 quotes
clear_thinking chat-template parameter default · 2 sources · 2 quotes
Deployment-time scalar effort control default · 1 source · 5 quotes
Cross-turn persistent reasoning history used · 3 sources · 3 quotes
Inference-time reasoning budget control used · 3 sources · 4 quotes
Reasoning parser used · 3 sources · 3 quotes
Generate-verify-refine loop used · 2 sources · 2 quotes
Medium-effort reasoning mode used · 2 sources · 2 quotes
Test-time compute scaling used · 2 sources · 2 quotes
Per-problem reasoning-budget control used · 1 source · 1 quote
Think-tag response formatting used · 1 source · 1 quote
Thinking modes (off and max) used · 1 source · 1 quote
Turn budget capping used · 1 source · 1 quote
Turn-aware prompting used · 1 source · 1 quote
Disabling reasoning via chat-template configuration optional · 5 sources · 5 quotes
Adaptive reasoning optional · 1 source · 1 quote
Configurable thinking or reasoning mode optional · 1 source · 1 quote
reasoning-enabled inference toggle optional · 1 source · 1 quote
Step-by-step thinking mode optional · 1 source · 1 quote
Task-appropriate maximum output length optional · 1 source · 1 quote
Parallel-fewest-step sampling evaluated · 2 sources · 2 quotes
Parallel test-time compute scaling evaluated · 1 source · 1 quote
Reward-optimized preferred reasoning length evaluated · 1 source · 1 quote
Qwen3 soft thinking switch not used · 2 sources · 2 quotes
User-configurable thinking-effort control not used · 1 source · 1 quote

KV cache management 42

Compressed KV caching core · 2 sources · 2 quotes
SWA Bounded Replay core · 2 sources · 4 quotes
Block-based KV cache core · 1 source · 1 quote
Cross-layer KV-cache reuse core · 1 source · 1 quote
Customized heterogeneous KV-cache layout core · 1 source · 1 quote
Fine-grained prefix hashing core · 1 source · 1 quote
KDA-aware prefix-cache management core · 1 source · 2 quotes
Persistent per-dialogue-context KV caching core · 1 source · 1 quote
Shared-free-list cache allocation core · 1 source · 1 quote
Write-back external KV-cache policy core · 1 source · 1 quote
Chunked prefill default · 6 sources · 6 quotes
Automatic cache default · 1 source · 1 quote
Prefix caching used · 5 sources · 5 quotes
Prompt caching used · 4 sources · 4 quotes
Language-model-only serving mode used · 3 sources · 3 quotes
RadixCache used · 2 sources · 2 quotes
Disable prefix caching for benchmarking used · 1 source · 1 quote
Encoder SWA bounded replay used · 1 source · 1 quote
Hierarchical KV caching used · 1 source · 1 quote
KDA with prefill cache used · 1 source · 1 quote
Key-value reuse in global attention layers used · 1 source · 1 quote
KV-cache sharing used · 1 source · 1 quote
Least-recently-used eviction used · 1 source · 1 quote
On-disk KV-cache storage used · 1 source · 1 quote
Overlapped host-memory prefetching used · 1 source · 1 quote
Persistent KV-cache management used · 1 source · 1 quote
Pre-scheduled tile chunking used · 1 source · 1 quote
Request-level prefix cache used · 1 source · 1 quote
Stateless in-memory prefix caching used · 1 source · 1 quote
Periodic cache checkpointing optional · 3 sources · 3 quotes
Zero SWA caching optional · 2 sources · 2 quotes
8-bit Mamba cache quantization evaluated · 1 source · 1 quote
Mamba prefix caching in align mode evaluated · 1 source · 1 quote

inference quantization 64

Quantization core · 3 sources · 3 quotes
FP4 KV-cache quantization core · 2 sources · 5 quotes
Dynamic activation scaling core · 1 source · 1 quote
Mixed-FP8 layers in an NVFP4 recipe core · 1 source · 1 quote
Mobile-specialized quantization schema core · 1 source · 1 quote
NVFP4 quantization for routed-expert GEMMs core · 1 source · 1 quote
Offline weight-layout permutation core · 1 source · 1 quote
Per-tensor FP8 quantization core · 1 source · 1 quote
Selective retention of BF16 precision core · 1 source · 1 quote
FP8 default · 14 sources · 14 quotes
NVFP4 quantization default · 5 sources · 5 quotes
NVFP4 KV-cache quantization default · 1 source · 2 quotes
Post-RoPE KV-cache quantization default · 1 source · 1 quote
FP8 KV-cache quantization used · 13 sources · 15 quotes
MXFP4 weight quantization used · 3 sources · 3 quotes
Post-training quantization used · 3 sources · 3 quotes
Four-Over-Six used · 2 sources · 3 quotes
GGUF used · 2 sources · 2 quotes
MXFP8 activation quantization used · 2 sources · 2 quotes
SSM cache quantization used · 2 sources · 2 quotes
AWQ INT4 weight quantization (W4A16) used · 1 source · 1 quote
BF16 inference used · 1 source · 1 quote
Block-wise E4M3 FP8 weight quantization used · 1 source · 1 quote
Channel-wise quantization used · 1 source · 1 quote
Embedding and KV-cache quantization used · 1 source · 1 quote
Flex_AWQ_SSZ used · 1 source · 1 quote
FP16 multimodal projector used · 1 source · 1 quote
FP4 precision for the attention indexer used · 1 source · 1 quote
FP8 inference used · 1 source · 1 quote
ModelOpt FP4 quantization used · 1 source · 1 quote
MXFP4 tensor packing used · 1 source · 1 quote
MXFP8 block-scale regrouping at load used · 1 source · 1 quote
NVFP4 quantization with modelopt used · 1 source · 1 quote
NVIDIA ModelOpt quantization used · 1 source · 1 quote
Post-quantization of MoE model layers used · 1 source · 1 quote
QuaRot used · 1 source · 1 quote
Random Hadamard transform used · 1 source · 1 quote
SpinQuant R1 rotation used · 1 source · 1 quote
Static activation quantization used · 1 source · 1 quote
Targeted 2-bit quantization used · 1 source · 1 quote
W4A16 quantization used · 1 source · 1 quote
W4A8 quantization used · 1 source · 1 quote
NVFP4 ModelOpt re-quantization optional · 2 sources · 2 quotes
Fine-grained FP8 quantization optional · 1 source · 1 quote
IQ4_XS quantization optional · 1 source · 1 quote
Mobile quantization optional · 1 source · 1 quote
Q3_K_L quantization optional · 1 source · 1 quote
Q4_0 quantization optional · 1 source · 1 quote
Q4_K_S quantization optional · 1 source · 1 quote
FP8 E4M3 quantization evaluated · 2 sources · 2 quotes
Max-based scaling evaluated · 2 sources · 2 quotes
MSE-based scaling evaluated · 2 sources · 2 quotes
Empirical bits-per-element budget selection evaluated · 1 source · 1 quote
MSE calibration evaluated · 1 source · 1 quote
Low-bit quantization mentioned · 1 source · 1 quote
In-flight block-wise FP8 weight quantization not used · 1 source · 1 quote

serving parallelism 17

Tensor parallelism (degree 4) default · 2 sources · 2 quotes
DeepEP used · 4 sources · 4 quotes
Prefill-decode disaggregation used · 4 sources · 4 quotes
Attention data parallelism used · 3 sources · 4 quotes
Encoder-Prefill-Decode disaggregation used · 2 sources · 2 quotes
Tensor parallelism (degree 8) used · 2 sources · 2 quotes
Topology-aware NVLink domain placement used · 2 sources · 2 quotes
Data-parallel vision encoding used · 1 source · 1 quote
Tensor parallelism for MoE layers used · 1 source · 1 quote
TensorRT-LLM all-reduce backend used · 1 source · 1 quote
Expert parallelism optional · 1 source · 1 quote
Low-precision MoE combine optional · 1 source · 1 quote
Token migration for balanced expert placement unclear · 1 source · 1 quote

inference scheduling 20

Cross-group pinning of cache-hit blocks core · 1 source · 1 quote
Asynchronous scheduling used · 3 sources · 3 quotes
Envoy-based proxy with custom orchestrator used · 1 source · 1 quote
Host-side scheduling optimization used · 1 source · 1 quote
Max sequences tuning used · 1 source · 1 quote
Per-node hard admission constraint used · 1 source · 1 quote
Continuous batching optional · 1 source · 1 quote
Exacto routing optional · 1 source · 1 quote
Increasing batch size for inference evaluated · 1 source · 1 quote

inference kernel 50

Batch-invariant deterministic kernels core · 2 sources · 3 quotes
Kernel fusion core · 2 sources · 2 quotes
Fused AttnRes merge and RMSNorm kernel core · 1 source · 1 quote
Separate-stream shared-expert GEMM overlap core · 1 source · 1 quote
Sparse pinned-host offload default · 1 source · 1 quote
CUDA Graph used · 3 sources · 3 quotes
CUDA Graph capture size reduction used · 2 sources · 2 quotes
FlashAttention 3 used · 2 sources · 2 quotes
MoE-side chunking used · 2 sources · 2 quotes
TensorRT-LLM multi-head attention backend used · 2 sources · 2 quotes
Dynamic load balancing used · 1 source · 1 quote
Exp-free TopK kernel used · 1 source · 1 quote
Expert-optimized Triton kernels used · 1 source · 1 quote
FA4 sheared-bias attention kernel used · 1 source · 1 quote
FlashAttention used · 1 source · 1 quote
FlashKDA used · 1 source · 1 quote
Fused mHC kernels used · 1 source · 1 quote
Fused QSA kernel used · 1 source · 1 quote
Fused RoPE-attention-RoPE-cast kernel used · 1 source · 1 quote
Hidden-dimension split Combine kernel used · 1 source · 1 quote
Host Codegen used · 1 source · 1 quote
KDA algorithm–system co-design used · 1 source · 1 quote
Marlin NVFP4 kernels used · 1 source · 1 quote
Mega-mHC used · 1 source · 1 quote
Persistent kernel used · 1 source · 1 quote
ROCm AITER sparse MLA attention backend used · 1 source · 1 quote
Single-Pass mHC used · 1 source · 1 quote
Sparse Flash Attention used · 1 source · 1 quote
Triton MoE kernel used · 1 source · 1 quote
Two-phase forward used · 1 source · 1 quote
FlashAttention 4 optional · 1 source · 1 quote
HPC-Ops attention backend optional · 1 source · 1 quote
HPC-Ops fused MoE backend optional · 1 source · 1 quote
FP8 GEMM evaluated · 1 source · 1 quote
Avoiding split-K not used · 1 source · 1 quote

context management 18

Preserved thinking history mode core · 7 sources · 7 quotes
Discard-all context management used · 6 sources · 7 quotes
Context compaction used · 3 sources · 3 quotes
Context folding used · 2 sources · 2 quotes
Thinking context management for tool use used · 2 sources · 2 quotes
Context management method used · 1 source · 1 quote
Context management strategy used · 1 source · 1 quote
Discarding tool-call history used · 1 source · 1 quote
Hierarchical context management used · 1 source · 1 quote
Keep-recent-k used · 1 source · 1 quote
Memory compression used · 1 source · 1 quote
Preserve thinking optional · 2 sources · 2 quotes
Modality-specific deployment optional · 1 source · 1 quote
Discard-75% evaluated · 2 sources · 2 quotes
Trajectory summarization and rollout re-initiation evaluated · 2 sources · 2 quotes
Summary-based context compression evaluated · 1 source · 1 quote

agentic scaffolding 93

Function calling core · 3 sources · 3 quotes
Harmony format core · 3 sources · 3 quotes
Role-based instruction hierarchy core · 2 sources · 2 quotes
Assistant output channels core · 1 source · 1 quote
Chain-of-thought in the analysis channel core · 1 source · 1 quote
Developer message format core · 1 source · 1 quote
Function-calling format core · 1 source · 1 quote
Harmony channel annotations core · 1 source · 1 quote
Harmony tool-call message format core · 1 source · 1 quote
Sandbox snapshots core · 1 source · 1 quote
Sandbox state forking core · 1 source · 1 quote
Automatic tool choice used · 4 sources · 4 quotes
Python tool used · 4 sources · 4 quotes
Browser tool used · 2 sources · 2 quotes
Generic tool-call parser used · 2 sources · 2 quotes
Multi-agent collaboration used · 2 sources · 2 quotes
Visual Search Tool used · 2 sources · 2 quotes
XML-based tool-call schema with DSML token used · 2 sources · 2 quotes
Agent Team mode used · 1 source · 1 quote
AgentENV used · 1 source · 1 quote
Agentic search used · 1 source · 1 quote
App-server mode with adapted tool schemas used · 1 source · 1 quote
AppArmor and eBPF sandbox policies used · 1 source · 1 quote
Asynchronous teammate spawning used · 1 source · 1 quote
Bash commands for context retrieval used · 1 source · 1 quote
Bash computer-use agent used · 1 source · 1 quote
Browsing tool with domain filtering used · 1 source · 1 quote
Commentary-channel preambles used · 1 source · 1 quote
Concurrent subagent orchestration used · 1 source · 1 quote
Current-turn-only reasoning-mode detection used · 1 source · 1 quote
Developer-defined function schemas used · 1 source · 1 quote
Dynamic tool loading used · 1 source · 1 quote
Harmony history stop-token normalization used · 1 source · 1 quote
Hy v4 tool-call parser used · 1 source · 1 quote
Indexed parallel tool calls used · 1 source · 1 quote
JSON Schema response formats used · 1 source · 1 quote
Jupyter Notebook code interpreter used · 1 source · 1 quote
Lead-agent interruption of teammates used · 1 source · 1 quote
Mini-harnesses used · 1 source · 1 quote
Native task delegation used · 1 source · 1 quote
Programmatic TypeScript tool calling used · 1 source · 1 quote
Prompt-enforced tool-call format used · 1 source · 1 quote
Qwen3 XML tool-call parser used · 1 source · 1 quote
ReAct Toolbelt used · 1 source · 1 quote
Retrieval-Augmented Search used · 1 source · 1 quote
Runtime command filter used · 1 source · 1 quote
Sandbox infrastructure (DSec) used · 1 source · 1 quote
Sandboxed Python execution loop used · 1 source · 1 quote
Scrollable browser text window used · 1 source · 1 quote
Shared task board with revision checks used · 1 source · 1 quote
Single-Bash-tool scaffold interface used · 1 source · 1 quote
Stateful Python tool used · 1 source · 1 quote
Step3p5 tool-call parser used · 1 source · 1 quote
System message format used · 1 source · 1 quote
Thinking with tools used · 1 source · 2 quotes
Tool calls within the thinking process used · 1 source · 1 quote
Tool output message format used · 1 source · 1 quote
Tool-Integrated Reasoning used · 1 source · 1 quote
Tools section in the system message used · 1 source · 1 quote
Typed tool arguments used · 1 source · 1 quote
TypeScript-like function schema syntax used · 1 source · 1 quote
Vision in the loop used · 1 source · 1 quote
Web search extension used · 1 source · 1 quote
XML-tagged tool-call format used · 1 source · 1 quote
XTML chat template used · 1 source · 1 quote
Advisor strategy optional · 2 sources · 4 quotes
MCP tool configuration optional · 2 sources · 2 quotes
Checkpointing optional · 1 source · 1 quote
Claude Code skills for Tinker optional · 1 source · 1 quote
Compact TXT image-path notation optional · 1 source · 1 quote
JSON mode optional · 1 source · 1 quote
OpenAI-style JSON content blocks optional · 1 source · 1 quote
Qwen3 Coder tool-call parser optional · 1 source · 1 quote
StreamableParser optional · 1 source · 1 quote
Tool calling optional · 1 source · 1 quote
Video understanding and editing optional · 1 source · 1 quote
Vision-driven UI coding optional · 1 source · 1 quote
Agentic tool calling mentioned · 1 source · 1 quote
Task-specification prompting mentioned · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modelfiled heredecoding strategyreasoning controlKV cache managementinference quantizationserving parallelisminference schedulinginference kernelcontext managementagentic scaffolding
GLM-5.3-Flash core—Multi-Token Prediction used—Always-on thinking mode coreConfigurable reasoning effort defaultReasoning-effort resolution and system-prompt injection defaultclear_thinking chat-template parameter optional——FP8 defaultMXFP8 block-scale regrouping at load usedFP8 KV-cache quantization optional—Identical cache-layout pinning across prefill and decode pools coreEncoder-Prefill-Decode disaggregation usedPrefill-decode disaggregation usedRound-robin routing for prefill-decode disaggregation used——ROCm AITER sparse MLA attention backend used——Video understanding and editing optionalVision-driven UI coding optional—
DeepSeek-V4.1-Flash core—DSpark coreThroughput-aware dynamic verification-length scheduling coreSpeculative decoding used—Capped linear reasoning-token length deduction coreConfigurable reasoning effort coreEffort-dependent exponential token-penalty schedule coreDeployment-time scalar effort control defaultReward-optimized preferred reasoning length evaluatedTest-time compute scaling evaluated—Compressed KV caching coreCross-layer KV-cache reuse coreSWA Bounded Replay coreEncoder SWA bounded replay usedLeast-recently-used eviction usedPersistent KV-cache management usedExact SWA KV reconstruction via full multi-layer replay not usedZero SWA caching not used—FP4 KV-cache quantization corePost-RoPE KV-cache quantization default—Encoder-Prefill-Decode disaggregation used—Latency-sensitive execution class with priority isolation usedPer-node hard admission constraint usedSub-NUMA partitioning with per-VM NUMA binding used—Kernel fusion coreFused RoPE-attention-RoPE-cast kernel usedMega-mHC usedSingle-Pass mHC used——Agent Team mode usedApp-server mode with adapted tool schemas usedAppArmor and eBPF sandbox policies usedAsynchronous teammate spawning usedFresh and fork teammate initialization modes usedLead-agent interruption of teammates usedMandatory per-turn tool calls with submission marker usedMulti-agent collaboration usedNative task delegation usedProgrammatic TypeScript tool calling usedShared task board with revision checks usedSingle-Bash-tool scaffold interface usedWeb search extension used—
Hy4-preview defaultcommunication optimization used—Multi-Token Prediction usedNEXTN speculative decoding usedSpeculative decoding used—Chain-of-thought reasoning defaultConfigurable reasoning effort used——FP8 usedLarge-model compression using quantization and speculative sampling optional—Tensor parallelism (degree 8) used——Kernel fusion used——Automatic tool choice usedHy v4 tool-call parser used—
DeepSeek-V4-Flash-0731 used—DSpark usedChat Prefix Completion optionalSpeculative decoding mentioned—Configurable reasoning effort used———————Function calling optionalJSON mode optional—
NVIDIA-Nemotron-3-Ultra-550B-A55B core—Recursive shared MTP-head drafting coreSpeculative decoding defaultBest-of-N scaffolding usedEAGLE usedMulti-Token Prediction used—Chain-of-thought reasoning usedConfigurable reasoning effort usedGenerate-verify-refine loop usedMedium-effort reasoning mode usedTurn budget capping usedTurn-aware prompting usedInference-time reasoning budget control optional—Chunked prefill defaultPrefix caching used8-bit Mamba cache quantization evaluatedPeriodic cache checkpointing evaluated—Mixed-FP8 layers in an NVFP4 recipe coreNVFP4 quantization for routed-expert GEMMs corePer-tensor FP8 quantization coreSelective retention of BF16 precision coreNVFP4 KV-cache quantization defaultFour-Over-Six usedFP16 cache storage with stochastic rounding usedFP8 usedFP8 KV-cache quantization usedHeuristic mixed per-layer precision quantization usedPost-training quantization usedRandom Hadamard transform usedSSM cache quantization usedW4A16 quantization usedBlock-scaled INT8 quantization with stochastic rounding evaluatedEmpirical bits-per-element budget selection evaluatedFP8 E4M3 quantization evaluatedMax-based scaling evaluatedMSE calibration evaluatedMSE-based scaling evaluated—Prefill-decode disaggregation usedTopology-aware NVLink domain placement usedAttention data parallelism optionalLow-precision MoE combine optional—Composite-key sorting for topology-aware GPU rank assignment usedNUMA binding of workers to GPU-local CPU sockets usedRegistering NVLink domains as Ray custom resources used—Marlin NVFP4 kernels usedMoE-side chunking used—Discard-all context management usedSummary-based context compression evaluated—Runtime command filter usedTool-Integrated Reasoning used—
MiMo-V2.6-Flash core—DFlash coreDistribution-matched draft-model fine-tuning usedThroughput-based draft-block sizing usedEAGLE optionalSpeculative decoding optional——Persistent per-dialogue-context KV caching coreAsynchronous cache offloading and restoration usedHierarchical KV caching used—Dynamic activation scaling coreFP8 low-precision speculative draft computation used—DeepEP used—Greedy feasible-rank placement by remaining capacity used———Mini-harnesses usedRequest timeouts and retries for agent loops usedMulti-agent collaboration mentioned—
DeepSeek-V4-Flash core——Cross-turn persistent reasoning history usedThink-tag response formatting usedThinking mode selection optional—Customized heterogeneous KV-cache layout coreDynamic allocation of fixed-size state-cache pools usedKV-cache and sparse-attention-kernel co-design usedOn-disk KV-cache storage usedPeriodic cache checkpointing optionalZero SWA caching optional—FP4 precision for the attention indexer usedFP8 KV-cache quantization used———Batch-invariant deterministic kernels coreDeepGEMM-based batch-invariant matrix multiplication usedDistributed shared memory for cross-SM attention data exchange usedDual-kernel batch-invariant attention decoding usedFused mHC kernels usedHost Codegen usedPer-SM accumulation buffers with deterministic global summation usedSeparate split-K outputs with deterministic reduction usedToken-order preprocessing and cross-rank buffer isolation for deterministic MoE backward usedAvoiding split-K not used——Agentic search usedRetrieval-Augmented Search usedXML-based tool-call schema with DSML token used—
MiMo-V2.5 used—Multi-Token Prediction usedSpeculative decoding usedEAGLE optional——Chunked prefill usedRequest-level prefix cache usedRadixCache not used—Block-wise E4M3 FP8 weight quantization usedFP8 used—DeepEP usedAttention data parallelism optional——FlashAttention 3 used—Discarding tool-call history usedMemory compression used—Bash commands for context retrieval used—
GLM-5.3 default—Multi-Token Prediction used—clear_thinking chat-template parameter defaultConfigurable reasoning effort default—Disable prefix caching for benchmarking used—FP8 defaultFP8 KV-cache quantization usedNVFP4 ModelOpt re-quantization optional——Max sequences tuning used——Context management strategy used——
Hy3 core—Speculative decoding coreEAGLE optionalSpeculative sampling optional—Configurable reasoning effort defaultReasoning parser usedChain-of-thought reasoning optional——FP8 usedFP8 KV-cache quantization optionalQuantization optionalLow-bit quantization mentioned—TensorRT-LLM all-reduce backend used——Triton MoE kernel usedHPC-Ops attention backend optionalHPC-Ops fused MoE backend optional——Generic tool-call parser used—
GLM-5.2 default—Multi-Token Prediction usedSpeculative decoding used—Configurable reasoning effort defaultDefault thinking mode usedMaximum thinking effort optional—Prefix caching usedRadixCache used—Flex_AWQ_SSZ usedFP8 inference usedQuaRot usedW4A8 quantization used—Attention data parallelism usedPrefill-decode disaggregation used—Asynchronous scheduling used—Sparse Flash Attention used—Discard-all context management usedHierarchical context management usedKeep-recent-k used——
MiniMax-M3 core——Test-time compute scaling usedAdaptive reasoning optionalConfigurable reasoning effort optionalThinking mode selection optional—Block-based KV cache coreAutomatic cache defaultPre-scheduled tile chunking used———Host-side scheduling optimization used—CUDA Graph usedDynamic load balancing usedExp-free TopK kernel usedFlashAttention usedPersistent kernel usedTwo-phase forward usedFP8 GEMM evaluated——Producer–Verifier adversarial harness loop usedReAct Toolbelt used—
DeepSeek-V3.2 used—Longest-trace selection usedMulti-stage candidate filtering usedTop-p and top-k sampling used—Generate-verify-refine loop usedreasoning-enabled inference toggle optionalDefault thinking mode evaluatedParallel test-time compute scaling evaluatedParallel-fewest-step sampling evaluatedSerial test-time compute scaling through context management evaluated——————Context management method usedTest-time context management for extending token budgets usedThinking context management for tool use usedDiscard-75% evaluatedDiscard-all context management evaluatedTrajectory summarization and rollout re-initiation evaluated—Jupyter Notebook code interpreter usedPrompt-based cold start for reasoning in tool use usedPrompt-enforced tool-call format usedThinking with tools usedTool calls within the thinking process used—
DeepSeek-V4-Flash-Vision-Exp default—Same-checkpoint target and draft weights defaultChat Prefix Completion optionalDSpark optional—Configurable reasoning effort optional——FP8 KV-cache quantization optional—————Compact TXT image-path notation optionalFunction calling optionalJSON mode optionalOpenAI-style JSON content blocks optional—
DeepSeek-V4-Pro core——Configurable reasoning effort usedCross-turn persistent reasoning history usedThink-tag response formatting usedThinking mode selection optional—Customized heterogeneous KV-cache layout coreDynamic allocation of fixed-size state-cache pools usedKV-cache and sparse-attention-kernel co-design usedOn-disk KV-cache storage usedPeriodic cache checkpointing optionalZero SWA caching optional—FP4 precision for the attention indexer usedFP8 KV-cache quantization used———Batch-invariant deterministic kernels coreDeepGEMM-based batch-invariant matrix multiplication usedDistributed shared memory for cross-SM attention data exchange usedDual-kernel batch-invariant attention decoding usedFused mHC kernels usedHost Codegen usedPer-SM accumulation buffers with deterministic global summation usedSeparate split-K outputs with deterministic reduction usedToken-order preprocessing and cross-rank buffer isolation for deterministic MoE backward usedAvoiding split-K not used——Agentic search usedRetrieval-Augmented Search usedSandbox infrastructure (DSec) usedXML-based tool-call schema with DSML token used—
DeepSeek-V4-Pro-0813 optional—Chat Prefix Completion optionalDSpark optional—Configurable reasoning effort optional———————Function calling optionalJSON mode optional—
Gemma 4 31B core—Multi-Token Prediction coreSpeculative decoding coreStandardized sampling configuration defaultKV-cache sharing between drafter and target usedTop-k over token clusters used—Control-token-enabled thinking mode coreDefault thinking mode coreConfigurable thinking or reasoning mode optionalStep-by-step thinking mode optional—Key-value reuse in global attention layers usedKV-cache sharing usedPrompt caching usedStateless in-memory prefix caching used—Mobile-specialized quantization schema coreQuantization coreChannel-wise quantization usedEmbedding and KV-cache quantization usedStatic activation quantization usedTargeted 2-bit quantization usedMobile quantization optionalQ4_0 quantization optionalPost-training quantization not used——Exacto routing optionalIncreasing batch size for inference evaluated——Excluding prior thinking from conversation history defaultModality-specific deployment optional—Function calling coreSandboxed Python execution loop used—
Inkling core—Multi-Token Prediction core—Configurable reasoning effort usedControllable thinking effort via system message and per-token cost used—Prompt caching used—NVFP4 quantization default—Fused reduce-scatter/all-gather collectives used——FA4 sheared-bias attention kernel used——Python tool usedClaude Code skills for Tinker optional—
Kimi K3 core—Fused recurrent replay kernel coreEAGLE-3-style draft-model fine-tuning used—Always-on thinking mode coreChain-of-thought reasoning defaultConfigurable reasoning effort defaultMaximum thinking effort defaultNatural-language reasoning-effort option message usedPer-problem reasoning-budget control usedstage-wise curriculum over reasoning-effort budget multiplier usedThinking mode selection via generation prefix used—Fine-grained prefix hashing coreKDA-aware prefix-cache management coreProjected-input caching for speculative KDA rollback coreShared-free-list cache allocation coreSparse hash-aligned KDA recurrent-state checkpoints coreUnified paged cache layout for KDA states and MLA KV coreWrite-back external KV-cache policy coreKDA with prefill cache usedKV-cache-aware placement of one-shot option messages used—Offline weight-layout permutation coreMXFP4 weight quantization usedMXFP8 activation quantization used—Zero-copy fused token permutation and unpermutation coreToken migration for balanced expert placement unclear—Cross-group pinning of cache-hit blocks coreDual-cluster consistent-hash failover for cache affinity corePrefix-cache-aware session affinity scheduling coreRequest-class resource-budget admission control coreRuntime-signal-based rollout concurrency auto-throttling coreWorkload-aware routed-expert GEMM scheduling core—Fused AttnRes merge and RMSNorm kernel coreFused latent down-projection and MoE-router GEMM coreGPU planning kernel for online expert placement coreSeparate-stream shared-expert GEMM overlap coreSide-stream overlap for inter-block attention-residual computation coreSynchronization-free static-shape MoE execution coreWarpDecode token-centric routed-expert decoding kernel coreFlashKDA usedKDA algorithm–system co-design used—Preserved thinking history mode coreContext compaction used—Natural-language in-context option instructions coreSandbox snapshots coreSandbox state forking coreAgentENV usedConcurrent subagent orchestration usedCopy-on-write memory and page-cache optimization usedDynamic tool loading usedIndexed parallel tool calls usedOverlayBD shared-image sandbox launch stack usedTyped tool arguments usedVision in the loop usedXTML chat template usedAgentic tool calling mentioned—
Laguna-S-2.1 core—DFlash optional—Interleaved thinking between tool calls coreConfigurable reasoning effort defaultMaximum thinking effort defaultThinking mode with automatic test-time compute budget defaultCross-turn persistent reasoning history usedThinking modes (off and max) usedUser-configurable thinking-effort control not used—Inference-side KV-cache reset on weight synchronization used—AWQ INT4 weight quantization (W4A16) usedFP8 KV-cache quantization usedFP8 W8A8 quantization with dynamic activation scaling usedMixed-precision INT4/INT8 layer-wise quantization usedPost-quantization of MoE model layers usedSpinQuant R1 rotation usedTRT-LLM inference optimization to NVFP4 on Blackwell usedIn-flight block-wise FP8 weight quantization not used——Envoy-based proxy with custom orchestrator used——Preserved thinking history mode used—Current-turn-only reasoning-mode detection usedStreaming-delta handling across block boundaries usedXML-tagged tool-call format usedCheckpointing optionalTask-specification prompting mentioned—
MiMo-V2.5-Pro used—Multi-Token Prediction usedSpeculative decoding usedEAGLE optionalMulti-layer EAGLE optional————DeepEP optional—————
MiMo-V2.6-Pro core—DFlash coreDistribution-matched draft-model fine-tuning usedThroughput-based draft-block sizing used—Chain-of-thought reasoning used—Persistent per-dialogue-context KV caching coreAsynchronous cache offloading and restoration usedHierarchical KV caching used—Dynamic activation scaling coreFP8 low-precision speculative draft computation used—DeepEP used—Greedy feasible-rank placement by remaining capacity used———Mini-harnesses usedRequest timeouts and retries for agent loops used—
NVIDIA-Nemotron-3.5-Lightning-30B-A3B core—Speculative decoding coreBest-of-N scaffolding usedMulti-Token Prediction usedDFlash optionalDSpark optionalConcurrency-aware draft-length tuning evaluated—Inference-time reasoning budget control used——Post-training quantization usedNVFP4 quantization optional—————Frontier-planning to execution-model routing coreBash computer-use agent used—
Qwen3.5-397B-A17B default—MTP-1 speculative decoding usedPresence Penalty usedMulti-Token Prediction optionalNEXTN speculative decoding optionalTask-specific sampling parameters optional—Default thinking mode defaultDisabling reasoning via chat-template configuration optionalTask- and mode-specific sampling parameter recommendations optionalTask-appropriate maximum output length optionalQwen3 soft thinking switch not used—Chunked prefill usedLanguage-model-only serving mode usedPrompt caching usedPrefix caching optional—FP8 usedNVFP4 quantization used—Data-parallel vision encoding usedTensor parallelism (degree 8) optional———Excluding prior thinking from conversation history defaultContext folding usedDiscard-all context management evaluated—Talker system prompt for voice characteristics usedMCP tool configuration optionalQwen3 Coder tool-call parser optionalTool calling optional—
Qwen3.6-35B-A3B default—Multi-Token Prediction optionalPresence Penalty optionalTask-specific sampling parameters optional—Default thinking mode defaultInterleaved thinking between tool calls defaultThinking mode selection defaultCross-turn persistent reasoning history usedDisabling reasoning via chat-template configuration optional—Chunked prefill usedPrefix caching usedPrompt caching usedLanguage-model-only serving mode optionalMamba prefix caching in align mode evaluated—FP8 KV-cache quantization usedFP8 optionalNVFP4 ModelOpt re-quantization optionalNVFP4 quantization optional——Asynchronous scheduling used—CUDA Graph capture size reduction used—Preserved thinking history mode usedPreserve thinking optional—Automatic tool choice usedQwen3 XML tool-call parser usedMCP tool configuration optional—
Qwen3.8-Flash-Next core—NEXTN speculative decoding usedMulti-Token Prediction optional—Default thinking mode default—Overlapped host-memory prefetching used—Fine-grained FP8 quantization optional—Tensor parallelism (degree 4) defaultExpert parallelism optional—Continuous batching optional—Shape-aware kernel dispatch for Hyper-Connection coreSparse pinned-host offload defaultCUDA Graph usedFused QSA kernel usedFusion of activation, gating, and reduction into GEMM epilogues usedHidden-dimension split Combine kernel usedSplit-K CuTe GEMM for low-batch Hyper-Connection Mix used—Preserved thinking history mode default——
Step-3.7-Flash used—EAGLE optionalMulti-layer EAGLE optionalSpeculative decoding optional—Reasoning parser usedConfigurable reasoning effort optional——FP16 multimodal projector usedFP8 usedFP8 KV-cache quantization usedGGUF usedModelOpt FP4 quantization usedNVFP4 quantization usedNVFP4 quantization with modelopt usedNVIDIA ModelOpt quantization usedIQ4_XS quantization optionalQ3_K_L quantization optionalQ4_K_S quantization optional——Asynchronous scheduling used—TensorRT-LLM multi-head attention backend usedFlashAttention 4 optional——Automatic tool choice usedGeneric tool-call parser usedPython tool usedStep3p5 tool-call parser usedVisual Search Tool usedAdvisor strategy optional—
gpt-oss-120b core—Grammar-constrained decoding usedBest-of-N scaffolding evaluated—Chain-of-thought reasoning coreConfigurable reasoning effort used——BF16 inference usedMXFP4 tensor packing usedMXFP4 weight quantization used—Tensor parallelism for MoE layers used——CUDA Graph usedExpert-optimized Triton kernels usedOptimized Triton MoE kernel with MXFP4 support used—Excluding prior thinking from conversation history used—Assistant output channels coreChain-of-thought in the analysis channel coreDeveloper message format coreFunction-calling format coreHarmony channel annotations coreHarmony format coreHarmony tool-call message format coreRole-based instruction hierarchy coreBrowser tool usedBrowsing tool with domain filtering usedCommentary-channel preambles usedDeveloper-defined function schemas usedHarmony history stop-token normalization usedInterleaving tool calls with chain-of-thought usedJSON Schema response formats usedPython tool usedRetaining chain-of-thought across tool-call turns usedScrollable browser text window usedStateful Python tool usedStateless Python tool reference implementation usedSystem message format usedTool output message format usedTools section in the system message usedTypeScript-like function schema syntax usedStreamableParser optional—