Model techniques map

1 733 methods · 1 502 filed · 2 955 quotes

Every method, filed where it belongs.

1 733 methods extracted from 164 documents, nested the way the curated taxonomy files them. Within each group, the strongest treatment any source gives a method comes first.

model architecture 312

filed at the root 7

Causal Encoder-Decoder (CED) architecture · Causal Encoder-Decoder architecture, Causal Encoder-Decoder (CED) core · 4 sources · 7 quotes
Per-Layer Embeddings (PLE) · per-layer embeddings core · 4 sources · 4 quotes
Asymmetric architecture core · 1 source · 1 quote
Encoder-free decoder-only Transformer architecture · encoder-free architecture core · 1 source · 1 quote
Regular Transformer decoder core · 1 source · 1 quote
Dense model architecture · 31B Dense used · 1 source · 1 quote

token mixer 123

Cross-layer attention sharing · attention sharing method core · 1 source · 1 quote
Micro-batch-level shared-state lifetime management · Micro-batch-level shared-state management used · 1 source · 1 quote
Dense gated attention with full RoPE and full gating · Dense GA, full RoPE, full gating evaluated · 1 source · 1 quote

softmax attention 32

Attention layers core · 1 source · 1 quote
DFlash attention core · 1 source · 1 quote
Softplus attention gating core · 1 source · 1 quote
Learned softmax denominator bias used · 1 source · 1 quote
Cascade attention not used · 2 sources · 2 quotes

global attention 4

Global attention · Global Attention (GA) core · 4 sources · 4 quotes
Unified Keys and Values · unified KV for global attention layers core · 2 sources · 2 quotes
Full attention used · 2 sources · 2 quotes
Multi-Head Attention (MHA) · Multi-Head Attention (MHA) for reliability used · 1 source · 1 quote

sliding window attention 12

Sliding Window Attention · SWA, Sliding Window Attention (SWA) core · 8 sources · 10 quotes
Hybrid Sliding Window Attention · Hybrid window attention core · 4 sources · 5 quotes
Sink-Augmented Sliding Window Attention · sink-augmented SWA core · 1 source · 1 quote
512-token Sliding Window Attention · + SWA-512 used · 2 sources · 2 quotes
Decoder SWA Bounded Replay · decoder bounded replay during prefill used · 1 source · 2 quotes
Dynamic Attention Window Size Training used · 1 source · 1 quote
Pure Sliding Window Attention used · 1 source · 1 quote
SWA-128 · Sliding Window Attention with size 128 used · 1 source · 1 quote
Fixed Sliding Window Attention · sliding-window baseline evaluated · 1 source · 1 quote
Per-Head Gating · Per-head gating with θswa=10,000, + per-head gating, θswa = 10,000 evaluated · 1 source · 1 quote
SWA-1024 · + SWA-1024 (interleaved 3:1) evaluated · 1 source · 1 quote

grouped-query attention 4

Grouped-query attention · GQA, GQA, 8 KV heads core · 11 sources · 13 quotes
Block-sparse grouped-query attention · block-sparse GQA core · 1 source · 1 quote
Multi-query attention · Multi-Query Attention (MQA), Multi-Query Attention (MQA) mode of MLA used · 2 sources · 2 quotes
Softplus-based per-head gating · Softplus-based per-head attention gating used · 1 source · 1 quote

multi-head latent attention 4

Gated Multi-Head Latent Attention · Gated MLA core · 3 sources · 3 quotes
MQA Mode of Multi-Head Latent Attention · MQA mode of MLA core · 2 sources · 2 quotes
Multi-Head Latent Attention · MLA, Multi-latent Attention core · 2 sources · 3 quotes
Multi-Head Latent Attention MHA/MQA Modes · MHA and MQA modes of MLA used · 1 source · 1 quote

attention sink 3

Learnable attention sink bias · Attention sink bias, learnable attention sink used · 5 sources · 5 quotes
Attention sink used · 1 source · 1 quote
RoPE with attention sink · RoPE + attention sink unclear · 1 source · 1 quote

sparse attention 57

DeepSeek Sparse Attention · DSA, DSA (DeepSeek Sparse Attention) core · 15 sources · 17 quotes
Compressed Sparse Attention · CSA, Compressed Sparse Attention (CSA) core · 6 sources · 6 quotes
Qwen Sparse Attention · Qwen Sparse Attention (QSA) core · 6 sources · 6 quotes
Sparse attention · sparse attention layers core · 6 sources · 6 quotes
Heavily Compressed Attention · Heavily Compressed Attention (HCA) core · 5 sources · 5 quotes
Gated DeepSeek Sparse Attention core · 3 sources · 3 quotes
Fixed-budget sparse-attention selection · fixed budget of 512 blocks or 2048 tokens core · 1 source · 1 quote
NoPE sparse multi-head latent attention · NoPE sparse MLA layers core · 1 source · 1 quote
QSA micro-block compression at ratio 4 core · 1 source · 1 quote
Sequential block processing core · 1 source · 1 quote
Token-wise compression core · 1 source · 1 quote
KV-outer sparse attention · KV outer gather Q used · 2 sources · 2 quotes
Sparse-attention continued pre-training with joint model and indexer optimization · Sparse Training Stage used · 2 sources · 2 quotes
CSA2 Full Mode · Full Mode used · 1 source · 2 quotes
Joint backbone and indexer training under sparse attention · sparse training used · 1 source · 1 quote
Natively trained sparsity used · 1 source · 1 quote
Sparse softmax attention · sparse softmax attention paradigm used · 1 source · 1 quote
Two-stage introduction of sparse attention · two-stage training method used · 1 source · 1 quote
Two-stage sparse attention · two-stage sparse-attention formulation used · 1 source · 1 quote
Sparse retrieval over long contexts evaluated · 1 source · 1 quote
Native Sparse Attention · NSA not used · 1 source · 1 quote

block selection 9

MiniMax Sparse Attention · MSA (MiniMax Sparse Attention), MiniMax Sparse Attention (MSA) core · 6 sources · 6 quotes
Block-level selection core · 1 source · 1 quote
Dynamic sparse selection core · 1 source · 1 quote
Main Branch core · 1 source · 2 quotes
Top-512 block selection · QSA keeps the best 512 blocks core · 1 source · 1 quote
Top-k block selection · Top- block selection, top-k selection used · 2 sources · 2 quotes
Block Max Pooling · Block Max Pool used · 1 source · 1 quote
Top-scoring compressed-block selection used · 1 source · 1 quote
Mixture of Block Attention · MoBA evaluated · 1 source · 1 quote

sparse attention indexer 22

IndexCache core · 3 sources · 3 quotes
IndexShare core · 3 sources · 3 quotes
Lightning Indexer core · 3 sources · 3 quotes
Compressed Sparse Attention 2 · Compressed Sparse Attention 2 (CSA2), cross-layer KV-cache and index reuse core · 2 sources · 6 quotes
Fine-grained token selection · fine-grained token selection mechanism core · 2 sources · 2 quotes
Hierarchical Sparse Indexer · Hierarchical sparse indexing core · 2 sources · 5 quotes
Compressed lightweight indexer core · 1 source · 1 quote
Index Branch core · 1 source · 2 quotes
IndexPool core · 1 source · 1 quote
MQA indexer core · 1 source · 1 quote
Reindex Mode · CSA2 Reindex Mode core · 1 source · 3 quotes
Reuse Mode · Cross-layer Top-K index reuse, CSA2 Reuse Mode core · 1 source · 3 quotes
Dense Warm-up Stage used · 2 sources · 2 quotes
Average pooling used · 1 source · 1 quote
Block-causal scoring used · 1 source · 1 quote
Detached indexer-input optimization used · 1 source · 1 quote
Indexer Warmup · indexer warmup with full attention used · 1 source · 2 quotes
Single-head index key · single-head key for index branch, single-head K_idx used · 1 source · 1 quote
Index Branch output · Index Branch value head output not used · 1 source · 1 quote
Index Branch value head · index value head not used · 1 source · 1 quote

fixed-pattern sparse attention 2

Sparse Global Attention Anchors core · 1 source · 1 quote
Local Block · forced sink and local window, forced sink and fixed local selection used · 1 source · 2 quotes

linear attention & state space 16

Linear attention · linear attention mechanism core · 2 sources · 2 quotes
Lightning Attention not used · 1 source · 1 quote

Mamba 3

Mamba-2 · Mamba-2 state-space model core · 5 sources · 6 quotes
Mamba core · 2 sources · 2 quotes
Mamba-2 SSM cache optional · 1 source · 1 quote

gated delta network 11

Gated DeltaNet · Gated Delta Net (GDN) module, Gated Delta Network linear attention core · 13 sources · 13 quotes
Kimi Delta Attention · KDA linear attention layers, Kimi Delta Attention (KDA) core · 6 sources · 6 quotes
Gated DeltaNet–sparse MoE hybrid core · 2 sources · 2 quotes
KDA · Kimi Distributed Attention (KDA) core · 2 sources · 2 quotes
Hybrid linear attention · Hybrid Linear Attention Mechanism core · 1 source · 1 quote
Gated DeltaNet with bounded sigmoid output gate · Sigmoid output gate for Gated DeltaNet, bounded sigmoid gate used · 1 source · 1 quote
Kimi Delta Attention with input-dependent full-rank output gate · Input-dependent full-rank output gating, input-dependent full-rank projection used · 1 source · 1 quote
Kimi Delta Attention with lower-bounded log-decay · lower-bounded decay used · 1 source · 1 quote
SimpleGDN evaluated · 1 source · 1 quote

hybrid layer stacking 13

Hybrid Attention · Hybrid attention mechanism, Hybrid Sparse and Linear Attention core · 27 sources · 29 quotes
Hybrid Mamba-Transformer · hybrid Mamba‑Transformer MoE, Mamba-Transformer core · 4 sources · 6 quotes
Mamba-2, Mixture-of-Experts, and Selective Attention Hybrid · Hybrid Mixture-of-Experts architecture, Interleaved Mamba-2 and MoE layers core · 4 sources · 7 quotes
Hybrid Mamba-Attention core · 3 sources · 4 quotes
Gated DeltaNet and Full Attention · Hybrid Attention Architecture core · 1 source · 1 quote
Gated DeltaNet and Gated Attention core · 1 source · 2 quotes
Gated DeltaNet and Qwen Sparse Attention · Hybrid Attention with QSA core · 1 source · 1 quote
Hybrid sparse mixture-of-experts Transformer architecture · hybrid sparse MoE Transformer backbone core · 1 source · 1 quote
Search-Based Sliding-Window Attention Pattern · SWA Pattern (Search-Based) evaluated · 1 source · 1 quote
Dense Attention Fallback · dense attention as a per-layer fallback unclear · 1 source · 1 quote

channel mixer 68

Contextual gating for N-gram embedding injection · contextual gating mechanism used · 1 source · 1 quote

dense feed-forward network 6

SiTU-GLU · SiTU-GLU activation function, SiTU-GLU activation (soft-capped SwiGLU) core · 2 sources · 3 quotes
SwiGLU · gated SwiGLU activation, gated SwiGLU [9] activation function used · 3 sources · 3 quotes
SwiGLU clamping · SwiGLU activation with clamping, SwiGLU activation function with clamping used · 3 sources · 3 quotes
Dense feed-forward network · dense Feed-Forward Network (FFN) used · 1 source · 1 quote
SiTU · Sigmoid Tanh Unit (SiTU) used · 1 source · 1 quote
SwiGLU clipping · SwiGLU clipping for activation control, SwiGLU-clip not used · 1 source · 1 quote

mixture of experts 61

Mixture of Experts · 198B total params / 11B activated params, 552B-parameter MoE core · 88 sources · 101 quotes
Sparse expert activation · Active parameters (sparse activation), active parameters core · 8 sources · 8 quotes
DeepSeekMoE · DeepSeekMoE framework core · 2 sources · 2 quotes
Mixture-of-Experts layers · MoE layers core · 2 sources · 2 quotes
MoE with routed and shared experts · 256 routed experts and 1 shared expert, Routed experts with shared expert core · 2 sources · 2 quotes
Asymmetric input/output activation split · asymmetric split core · 1 source · 1 quote
Dispatch recomputation · recomputing dispatch core · 1 source · 1 quote
Gated DeltaNet MoE · gated-delta-networks MoE architecture, routed and shared experts in MoE core · 1 source · 2 quotes
Hybrid Mixture of Experts · Hybrid Attention Mixture-of-Experts (MoE) core · 1 source · 1 quote
MegaMoE used · 1 source · 1 quote
MoE with 256 experts and top-8 routing · Sparse MoE with 256 experts and 8 top-k used · 1 source · 1 quote
Routed expert output modulation · routed expert modulation used · 1 source · 1 quote

expert routing 15

Top-8 expert routing · 192 experts, top-8 activated, top-8 routed experts core · 5 sources · 5 quotes
Token-level expert routing · token routing through 8 of 288 experts, routes each token through 8 of 288 experts core · 3 sources · 3 quotes
Keep Routing core · 1 source · 1 quote
Routed experts · 192 routed experts (top-8) core · 1 source · 1 quote
Token-choice routing with softplus gating · Token-choice router with softplus gating core · 1 source · 1 quote
Top-6 expert routing · each token is routed to 6 of 256 experts core · 1 source · 1 quote
Fixed Top-k routing with frozen bias · fixed Top- k selection with a frozen bias default · 1 source · 1 quote
Anticipatory Routing used · 2 sources · 2 quotes
DP-aware routing · DP-aware routing mechanism used · 1 source · 1 quote
Hash routing used · 1 source · 1 quote
Latent-space routing used · 1 source · 1 quote
Token-choice routing · Token-choice mixture-of-experts routing used · 1 source · 1 quote
Top-4 expert routing · top-4 expert selection, top-4 routed expert selection used · 1 source · 2 quotes
Loss-spike-triggered Anticipatory Routing optional · 1 source · 1 quote

expert load balancing 15

Auxiliary-loss-free load balancing · auxiliary-loss-free strategy, Auxiliary-loss-free MoE load balancing core · 3 sources · 5 quotes
Quantile Balancing · QB load balancing for mixture-of-experts, QB for MoE load balancing core · 3 sources · 5 quotes
Modality-specific auxiliary-loss-free load balancing · modality-specific load balancing core · 1 source · 1 quote
Auxiliary-loss load balancing · auxiliary loss used · 1 source · 1 quote
Exact coordinate minimization used · 1 source · 1 quote
Expert bias update factor used · 1 source · 1 quote
Expert Parallelism Load Balancing · EPLB used · 1 source · 1 quote
formHC · formHC (MoE routing/load-balancing method) used · 1 source · 1 quote
Histogram-based quantile estimation · Histogram estimation, B uniform bins used · 1 source · 2 quotes
Load balancing used · 1 source · 1 quote
Persistent load balancing used · 1 source · 1 quote
Pooled-global-batch quantile estimation used · 1 source · 1 quote
Round-robin load balancing · --load-balance-method round_robin optional · 1 source · 1 quote
MaxVio · MaxVio metric evaluated · 1 source · 1 quote

shared experts 7

Shared Experts · Shared Expert, 1 shared expert core · 7 sources · 7 quotes
No Shared Experts · contains no shared experts, Sparse MoE FFNs without shared experts core · 3 sources · 3 quotes
10 Routed + 1 Shared Experts Activated per Token · 10 routed + 1 shared activated per token, 10 Routed + 1 Shared core · 2 sources · 2 quotes
Shared Experts Active on Every Token · 2 shared experts active on every token core · 2 sources · 2 quotes
MoE with Expert Sinks · Mixture-of-Experts with expert sinks core · 1 source · 1 quote

fine-grained experts 2

DeepSeekMoE shared and fine-grained routed experts · shared and fine-grained routed experts core · 1 source · 1 quote
Granular Mixture of Experts · Granular MoEs not used · 2 sources · 2 quotes

latent mixture of experts 5

LatentMoE · Latent Mixture of Experts, Latent Mixture-of-Experts (LatentMoE) core · 7 sources · 9 quotes
Normalized LatentMoE · Stable LatentMoE, Stable LatentMoE framework core · 6 sources · 7 quotes
Hybrid Latent Mixture of Experts core · 1 source · 1 quote
Latent-dimension reduction used · 1 source · 1 quote

positional encoding 14

Relative attention · Relative positional embedding core · 2 sources · 2 quotes
No Position Encoding · No Position Encoding on MLA layers, No Position Encoding (NoPE) core · 1 source · 2 quotes
Omitting RoPE in attention layers · do not use RoPE in attention layers core · 1 source · 1 quote
Per-layer-type rotary position scales · per-layer-type rotary scales core · 1 source · 1 quote
YaRN · Modifying the model configuration file, Static YaRN context length extension used · 12 sources · 19 quotes
Rotary Position Embedding · RoPE, Rotary Positional Embedding (RoPE) used · 5 sources · 5 quotes
2D rotary position embedding · 2D rotary position embeddings (2D-RoPE), 2D-RoPE used · 3 sources · 3 quotes
Proportional Rotary Position Embedding · Proportional RoPE (p-RoPE), proportional rotary position embeddings used · 2 sources · 2 quotes
2D coordinate-based positional embeddings used · 1 source · 1 quote
Partial RoPE used · 1 source · 1 quote
RoPE scaling · RoPE scaling for context extension, RoPE scaling techniques optional · 3 sources · 3 quotes
Gated attention with partial RoPE · Gated attention with partial RoPE (50%), + GA Partial RoPE (50%) evaluated · 1 source · 1 quote
Temporal-Modality Rotary Position Embedding · TM-RoPE not used · 1 source · 1 quote

normalization & residual 28

Manifold-Constrained Hyper-Connections · mHC core · 8 sources · 8 quotes
Attention Residuals (AttnRes) · Attention Residuals, AttnRes core · 7 sources · 7 quotes
Gated Residual · Gated Residual (GR), Gated residual connections core · 6 sources · 8 quotes
Identity Hyper-Connections · iHC (identity Hyper-Connections), Identity Hyper-Connections (iHC) core · 3 sources · 3 quotes
Hyper-Connections · Hyper-Connections (HC) core · 2 sources · 2 quotes
Per-Branch Scalar Write Gate · Per-branch scalar residual write gate, write gate core · 2 sources · 2 quotes
Single-Pass mHC · Single-PassmHC residual-stream mixing, Single-PassmHC core · 2 sources · 3 quotes
All-branch Elementwise Read Gate core · 1 source · 1 quote
Block AttnRes · block attention residual, Block Attention Residuals core · 1 source · 2 quotes
Elementwise Data-Dependent Residual Read Gate · read gate core · 1 source · 1 quote
Four-Branch Residual Stream · widens the residual stream into 4 branches core · 1 source · 1 quote
GatedNorm · GatedNorm elementwise self-gating, elementwise self-gate after RMSNorm core · 1 source · 3 quotes
Independent Per-Branch Normalization · group RMSNorm over the widened stream core · 1 source · 2 quotes
Low-Rank Gated Mix core · 1 source · 1 quote
Two-Phase Block AttnRes Schedule · two-phase Block AttnRes kernel schedule, two-phase schedule core · 1 source · 1 quote
Widened Residual Stream · widening the residual stream core · 1 source · 1 quote
RMSNorm · RMSNorm normalization, root mean square normalization (RMSNorm) used · 3 sources · 3 quotes
Bounded Positive Gates used · 1 source · 1 quote
Data-Dependent Residual Read and Write Operators · making and data-dependent used · 1 source · 1 quote
Post-Embedding RMSNorm used · 1 source · 1 quote
Pre-LN · Pre-LN (pre-normalization) placement, Pre-LN placement used · 1 source · 1 quote
Zero-Centered RMSNorm used · 1 source · 1 quote
Simplified AltUp · Simplified AltUp residual widening, a simplified variant of AltUp evaluated · 1 source · 1 quote
Full Attention Residuals mentioned · 1 source · 1 quote
QK-Clip · QK-Clip technique, QK clipping for activation control not used · 2 sources · 2 quotes
Sparse Gated Residual Writes · sparse writes not used · 1 source · 1 quote

prediction head 5

Multi-Token Prediction · 3.8B MTP layer, MTP core · 30 sources · 40 quotes
Multi-Token Prediction for residual codebooks · multi-token prediction (MTP) module used · 1 source · 1 quote
Nemotron Hybrid Multi-Token Prediction · nemotron_h_mtp optional · 1 source · 1 quote

multimodal architecture 57

Encoder-free multimodal architecture · Encoder-free decoder-only architecture, encoder-free architecture core · 6 sources · 6 quotes
Early fusion multimodal training · Early fusion training on multimodal tokens, Unified Vision-Language Foundation core · 5 sources · 5 quotes
Discrete token encoding for audio · discrete token encoding, discrete token encoding for audio inputs core · 4 sources · 4 quotes
Hierarchical patch encoder · Hierarchical patch encoder for images core · 4 sources · 4 quotes
Mixed-modality training · mixed modalities from the start core · 4 sources · 4 quotes
Native multimodal understanding · native multimodal in one model, native omnimodal architecture core · 3 sources · 3 quotes
3×3 pixel-unshuffle downsampling core · 2 sources · 2 quotes
Joint vision-language pre-training core · 2 sources · 2 quotes
Joint multimodal input and understanding core · 1 source · 1 quote
Language backbone core · 1 source · 1 quote
Multimodal mixture-of-experts Transformer · multimodal mixture of experts core · 1 source · 1 quote
Native multimodal encoding · natively multimodal core · 1 source · 1 quote
Native multimodal visual understanding core · 1 source · 1 quote
Native visual and audio understanding · multimodal understanding core · 1 source · 1 quote
Single shared multimodal backbone · single shared backbone core · 1 source · 1 quote
Thinker–Talker architecture core · 1 source · 1 quote
Vision encoder core · 1 source · 1 quote
Vision encoder and multimodal aligner · vision encoder and aligner core · 1 source · 1 quote
Vision encoder–MLP projector pathway · vision encoder and an MLP projector core · 1 source · 1 quote
Visual modules for multimodal understanding · incorporating visual modules core · 1 source · 1 quote
Raw audio projection into the LLM embedding space · raw audio chunk projection used · 2 sources · 2 quotes
Vision Transformer (ViT) · Vision Transformer, ViT vision encoder used · 2 sources · 2 quotes
2×2 pixel-shuffle downsampling used · 1 source · 1 quote
Audio encoder initialized from MiMo-Audio used · 1 source · 1 quote
Audio Transformer (AuT) · attention-encoder-decoder model AuT used · 1 source · 1 quote
Causal streaming ConvNet codec decoder used · 1 source · 1 quote
Data-parallel-first multimodal encoding · encoding runs data-parallel first used · 1 source · 1 quote
Decoupled Encoder Process (DEP) · decoupled encoder process used · 1 source · 1 quote
Dedicated multimodal encoders used · 1 source · 1 quote
Disaggregated encoder training · disaggregated encoder design used · 1 source · 1 quote
Dual-format coordinate supervision · coordinate supervision used · 1 source · 1 quote
Explicit text-string timestamps used · 1 source · 1 quote
Frozen encoders during pre-training used · 1 source · 1 quote
Joint full-model multimodal training used · 1 source · 1 quote
Lightweight vision embedding module used · 1 source · 1 quote
Linear-projection patch embedding used · 1 source · 1 quote
Mel-spectrogram audio front end used · 1 source · 1 quote
MoonViT-V2 · MoonViT-V2 vision encoder used · 1 source · 1 quote
Multimodal input · image input used · 1 source · 1 quote
Projector warmup used · 1 source · 1 quote
Residual Vector Quantization (RVQ) speech representation · RVQ-based speech representation used · 1 source · 1 quote
Reusing the image token for video frames · reusing image token for video frames used · 1 source · 1 quote
Single-matrix-multiplication vision projection · single matmul vision projection, single large matmul projection for vision used · 1 source · 1 quote
Staged vision-encoder freezing used · 1 source · 1 quote
Two-layer MLP vision projector · two-layer MLP projector used · 1 source · 1 quote
Two-stage text-then-multimodal pre-training · two-stage pre-training strategy used · 1 source · 1 quote
Universal Speech Model (USM)-based audio encoder · Universal Speech Model-based audio encoder, USM-based audio encoder used · 1 source · 1 quote
Variable aspect-ratio image handling · variable aspect ratios used · 1 source · 1 quote
Configurable visual token budget optional · 3 sources · 3 quotes
Video preprocessor longest-edge configuration optional · 1 source · 1 quote

context capacity 10

N-gram embedding lookup · Local-context N-gram embedding lookup, N-gram Embedding core · 6 sources · 6 quotes
1M-token context window · 1M context length, 1M context core · 4 sources · 4 quotes
N-gram embedding host-memory offload and prefetch · asynchronously offloaded to host memory, asynchronous prefetching core · 4 sources · 6 quotes
Engram · Engram conditional memory, Engram conditional memory module core · 3 sources · 5 quotes
256K context window core · 1 source · 1 quote
Single-layer N-gram embedding placement · A single N-gram embedding layer core · 1 source · 1 quote
Progressive context extension · context window extension, progressive context extension curriculum used · 2 sources · 2 quotes
Multi-head hashing used · 1 source · 1 quote

training objective 29

filed at the root 1

SigLIP sigmoid contrastive loss · sigmoid contrastive loss used · 1 source · 1 quote

language modelling objective 6

Cross-entropy objective · Cross-entropy training objective used · 1 source · 1 quote
Next-token prediction · next-token prediction objective used · 1 source · 1 quote
Sampled-token objective used · 1 source · 1 quote
Stricter token constraints during training used · 1 source · 1 quote
Text pre-training used · 1 source · 1 quote
Fill-in-the-middle (FIM) completion · fill-in-the-middle completion, FIM Completion optional · 1 source · 1 quote

multi-token prediction objective 6

Multi-Token Prediction · Multi-Token Prediction (MTP), Multiple-Token Prediction used · 3 sources · 5 quotes
Multi-Token Prediction Boosting · MTP Boosting, MTP-boosting phase used · 3 sources · 5 quotes
Multi-step MTP training · MTP: trained with multi-steps used · 2 sources · 2 quotes
Continued pretraining for MTP layers used · 1 source · 1 quote
Frozen-backbone MTP-head fine-tuning used · 1 source · 1 quote
Shared-weight Multi-Token Prediction · shared-weight MTP objective used · 1 source · 1 quote

distillation objective 7

Full-vocabulary logit distillation used · 2 sources · 2 quotes
Temperature-scaled forward KL distillation · temperature-scaled forward-KL loss used · 2 sources · 2 quotes
Dense-attention distillation · dense distillation used · 1 source · 1 quote
Hidden-state caching for KL distillation used · 1 source · 1 quote
Quantization-aware distillation · Quantization-aware distillation (QAD) used · 1 source · 1 quote
Reverse KL divergence · reverse KL divergence loss used · 1 source · 1 quote
Logit matching evaluated · 2 sources · 3 quotes

auxiliary loss 9

KL alignment loss · KL-divergence loss, KL divergence loss for indexer used · 3 sources · 4 quotes
Sample-level attention masking used · 3 sources · 3 quotes
Domain-specific KL regularization strength · Domain-adaptive KL regularization strength, varying strengths of KL regularization used · 1 source · 2 quotes
Gradient detachment · Gradient Detach, gradient detach for KL loss used · 1 source · 2 quotes
Language modeling loss plus KL loss · LM Loss + KL Loss used · 1 source · 1 quote
Loss masking · Loss masking of flagged trajectory content, mask used · 1 source · 1 quote
MoE sequence auxiliary loss · MoE sequence auxiliary loss coefficient used · 1 source · 1 quote
Sequence-level balance loss used · 1 source · 1 quote
Unfinished-trajectory loss masking · mask the loss on unfinished trajectories used · 1 source · 1 quote

optimization 155

filed at the root 4

Updated hyperparameter scaling law · refit the scaling law, dedicated scaling-law studies core · 3 sources · 4 quotes
Momentum update with Sinkhorn balancing core · 1 source · 1 quote
Fused loss computation used · 1 source · 1 quote
Phase-by-phase training cost accounting · phase-by-phase cost accounting used · 1 source · 1 quote

optimizer 29

Muon · Muon optimizer, Muon matrix-based optimizer core · 11 sources · 13 quotes
Category-specific assignment of Muon and AdamW · Tailored Training Recipe, division of labour between Muon and AdamW core · 3 sources · 4 quotes
Split fused gradients before orthogonalization · splitting of fused parameters core · 2 sources · 3 quotes
AdamW for the MoE router · we use AdamW for the router core · 1 source · 1 quote
Per-Head Muon · head-wise Muon, Per-Head Muon optimizer default · 4 sources · 5 quotes
Eight-step Newton–Schulz iteration · Eight-step Newton–Schulz orthogonalization, Newton–Schulz iteration to 8 steps default · 1 source · 1 quote
Sinkhorn-balanced update · Momentum-based Sinkhorn-balanced optimizer default · 1 source · 2 quotes
AdamW · AdamW optimizer used · 3 sources · 4 quotes
Nesterov momentum · Nesterov-accelerated momentum, Nesterov trick used · 2 sources · 2 quotes
AdamW for attention and GDN output gates used · 1 source · 1 quote
AdamW for input embeddings and output head used · 1 source · 1 quote
Asynchronous Micro-Group pipeline · an asynchronous Micro-Group pipeline used · 1 source · 1 quote
Canzona used · 1 source · 1 quote
CUDA graph capture of the optimizer step · We capture the whole step in a CUDA graph used · 1 source · 1 quote
Hybrid Newton–Schulz iterations used · 1 source · 1 quote
Hybrid optimization with Muon and Adam used · 1 source · 1 quote
Moonlight-style learning-rate scaling used · 1 source · 1 quote
Muon orthogonalization accuracy refinement · orthogonalization accuracy used · 1 source · 1 quote
Muon Split used · 1 source · 1 quote
Muown used · 1 source · 1 quote
Peer-to-peer shard retrieval for Muon orthogonalization · P2P communication used · 1 source · 1 quote
Polar Express coefficient schedule used · 1 source · 1 quote
Selected number of Newton–Schulz iterations · the number of iteration steps used · 1 source · 1 quote

learning-rate schedule 13

Cosine decay · Cosine decay learning rate schedule, cosine learning rate schedule default · 4 sources · 5 quotes
Warmup-Stable-Decay · Warmup Stable Decay learning rate schedule, Warmup Stable Decay (WSD) used · 4 sources · 5 quotes
Batch-size warmup · Ramping the batch size over early training, Batch-size warmup during early training used · 2 sources · 4 quotes
Training without batch-size warmup used · 2 sources · 2 quotes
Engram learning-rate scaling · learning rate of Engram is scaled by 5× used · 1 source · 1 quote
Fixed elevated learning rate used · 1 source · 1 quote
Learning-rate annealing used · 1 source · 1 quote
Linear warmup followed by cosine decay · linear warmup then cosine decay used · 1 source · 1 quote
Scaling-law fit · new scaling-law fit used · 1 source · 1 quote
Scheduled batch-size growth · batch size scheduling strategy used · 1 source · 1 quote
WSD-specific scaling law · WSD learning-rate scaling law used · 1 source · 1 quote

training precision 19

NVFP4 · NVFP4 4-bit Training Format, ultraefficient 4-bit NVFP4 training format core · 10 sources · 13 quotes
NVFP4 pre-training · NVFP4 pre-training recipe, Pre-training with NVFP4 quantization core · 5 sources · 5 quotes
BF16 mixed-precision training default · 1 source · 1 quote
BF16 · BF16 precision, BF16 testing used · 5 sources · 5 quotes
FP8 mixed-precision training · FP8 mixed precision used · 4 sources · 4 quotes
FP4+FP8 mixed precision · FP4 + FP8 Mixed used · 2 sources · 2 quotes
FP8-precision reinforcement learning · RL was done in FP8 precision used · 2 sources · 2 quotes
MXFP8 · keep these layers in MXFP8 used · 2 sources · 2 quotes
E2M1 · E2M1 Floating Point, E2M1 datatype used · 1 source · 1 quote
FP32 attention-output retention used · 1 source · 1 quote
FP32 gradient reduction used · 1 source · 1 quote
FP8 storage for the residual state · FP8 storage for the widened residual state, keep the residual state in FP8 used · 1 source · 1 quote
High-precision final network layers used · 1 source · 1 quote
Mixed-FP8 quantization · mixed-FP8 used · 1 source · 1 quote
Mixed-precision training · mixed-precision training/rollouts used · 1 source · 1 quote
NVFP4 fine-grained micro-block scaling used · 1 source · 1 quote
Two-dimensional block quantization used · 1 source · 1 quote
BF16 gradient reduction not used · 1 source · 1 quote

quantization-aware training 15

Quantization-Aware Training · Quantization-Aware Training (QAT) core · 8 sources · 8 quotes
Quantize-Dequantize Training · quantize–dequantize (QDQ) core · 1 source · 1 quote
FP4 Quantization · FP4 (MXFP4) quantization, MXFP4 quantization default · 4 sources · 4 quotes
FP4 Quantization-Aware Training · MXFP4 quantization-aware training, MXFP4 quantization-aware training (QAT) used · 3 sources · 3 quotes
MXFP4 Weights with MXFP8 Activations · MXFP4/MXFP8 quantization-aware training used · 2 sources · 3 quotes
Stochastic Rounding used · 2 sources · 3 quotes
INT4 Quantization-Aware Training used · 1 source · 1 quote
MXFP4 Quantization-Aware Post-Training used · 1 source · 1 quote
NVFP4 Training · quantized pretraining with NVFP4 used · 1 source · 1 quote
Per-Block Scalar Scaling used · 1 source · 1 quote
Q4_0 Quantization Format used · 1 source · 1 quote
Random Hadamard Transforms used · 1 source · 1 quote
Stochastic Rounding for Mamba Cache · mamba-cache-stochastic-rounding used · 1 source · 1 quote
Stochastic Rounding of Gradients · stochastic rounding on gradients used · 1 source · 1 quote
Straight-Through Estimator used · 1 source · 1 quote

training stability 17

Weight padding used · 2 sources · 2 quotes
Anti-hallucination training used · 1 source · 1 quote
Behavioral regularization used · 1 source · 1 quote
Cross-replica model-weight hash consistency checks · Cross-replica Hash Checks used · 1 source · 1 quote
Exponential moving average of checkpoints · exponential moving average (EMA) used · 1 source · 1 quote
Gradient clipping used · 1 source · 1 quote
Learning-rate elevation for training stability stress testing · raising the learning rate used · 1 source · 1 quote
Loss masking for excessively stale tokens used · 1 source · 1 quote
Off-policy sample filtering · Dropping off-policy and noisy samples used · 1 source · 1 quote
Per-token regularization for off-policy RL used · 1 source · 1 quote
Soft dropping used · 1 source · 1 quote
Training stability stress testing · stress tests used · 1 source · 1 quote
Weight clipping · weight-clipping mechanism used · 1 source · 1 quote
Activation or logit clipping not used · 1 source · 1 quote

training parallelism 29

Communication-Computation Overlap core · 1 source · 2 quotes
MoonEP · MoonEP expert placement planning core · 1 source · 4 quotes
Redundant-Expert Capacity Reservation core · 1 source · 1 quote
Sequence Parallelism for Activations · sequence parallelism (SP) for activations core · 1 source · 1 quote
Expert Parallelism · MoE Expert Parallelism (EP), expert parallel used · 10 sources · 13 quotes
Tensor Parallelism used · 5 sources · 6 quotes
Pipeline Parallelism used · 2 sources · 2 quotes
Cache-Based Pipeline Communication used · 1 source · 1 quote
Context Parallelism used · 1 source · 1 quote
Dynamic Context Parallelism for Large Multimodal Samples · Dynamic CP in multimodal encoder used · 1 source · 1 quote
Fully Balanced Expert-Parallel Training used · 1 source · 1 quote
GPU Planning Kernel for Redundant Expert Migration · online planning of redundant experts used · 1 source · 1 quote
Hybrid ZeRO Bucket Assignment for Muon used · 1 source · 1 quote
KDA Context Parallelism · KDA Context Parallelism (KCP) used · 1 source · 1 quote
Load-Balanced Image Sharding · Balanced image sharding used · 1 source · 1 quote
NS-FLOP-Balanced Static Parameter Partitioning · An -balanced static partitioner used · 1 source · 1 quote
Pipeline Payload Extensions used · 1 source · 1 quote
pipeline-bubble scheduling of ViT computation · scheduled into pipeline bubbles used · 1 source · 1 quote
SConv-Aware Tensor-Parallel Sharding · sconv-aware TP sharding used · 1 source · 1 quote
SM-Level Context Parallelism · intra-device context parallelism used · 1 source · 1 quote
Tensor and Expert Parallelism with Degree Eight · TEP8 used · 1 source · 1 quote
ZeRO-3 Optimizer State Sharding used · 1 source · 1 quote
Data-Weighted Data Parallelism · DWDP mentioned · 1 source · 1 quote

training runtime 29

Composable activation storage policies core · 1 source · 1 quote
In-flight recovery system core · 1 source · 1 quote
Offline checkpoint merging used · 3 sources · 3 quotes
Asynchronous checkpointing · Asynchronous background checkpointing used · 2 sources · 2 quotes
Asymmetric local-read and single-writer cache paths · Asymmetric read/write cache paths used · 1 source · 1 quote
Autotune configuration generation · autotune used · 1 source · 1 quote
Batch-level embedding prefetch · Embedding prefetch used · 1 source · 1 quote
Caching the distributed checkpoint save plan · cache the distributed save plan used · 1 source · 1 quote
CPU-resident optimizer states · Optimizer states stay in CPU memory used · 1 source · 1 quote
Cross-rank remote activation offloading · remotely offload activations used · 1 source · 1 quote
Element-wise activation recomputation used · 1 source · 1 quote
GPU-to-GPU weight synchronization over GPUDirect RDMA · GPU-to-GPU weight transfer used · 1 source · 1 quote
Heartbeat-driven fault tolerance used · 1 source · 1 quote
In-flight weight updates used · 1 source · 1 quote
Job-level eviction and reclaim · Per-job eviction and reclaim used · 1 source · 1 quote
Overlapping NCCL transfers with device-to-host copies · overlapped NCCL transfers with D2H copies used · 1 source · 1 quote
Persistent checkpoint worker processes used · 1 source · 1 quote
Persistent shared-storage cache for compiled artifacts · Persistent warm cache on shared storage used · 1 source · 1 quote
Post-iteration NVMe offloading of training states · offload training states ... to NVMe used · 1 source · 1 quote
Pre-admission hardware stress testing · pre-flight checks used · 1 source · 1 quote
Same-node sticky pod respawn · Sticky pod respawn used · 1 source · 1 quote
Seeding node-local storage from a warm shared cache · Node-local seeding at startup used · 1 source · 1 quote
Slice-Granularity Elasticity used · 1 source · 1 quote

data curation 160

filed at the root 5

Data curation used · 2 sources · 2 quotes
Training data deduplication and filtering · deduplication and filtering used · 2 sources · 2 quotes
Multilingual corpus expansion · expands its multilingual corpus used · 1 source · 1 quote
Nemotron-CC pipeline used · 1 source · 1 quote
Product co-design for training data used · 1 source · 1 quote

data sourcing 17

Expert co-created training data used · 2 sources · 2 quotes
GitHub Crawl · GitHub Crawl via REST and S3 APIs used · 2 sources · 2 quotes
Agentic knowledge graph construction used · 1 source · 1 quote
Common Crawl used · 1 source · 1 quote
EssentialWeb used · 1 source · 1 quote
FinePDFs used · 1 source · 1 quote
High-recall web-data curation · high-recall web data used · 1 source · 1 quote
Long-document curation used · 1 source · 1 quote
Nemotron-3-Ultra corpus used · 1 source · 1 quote
Nemotron-CC used · 1 source · 1 quote
Nemotron-Post-Training-v3 used · 1 source · 1 quote
Streaming training-data ingestion · stream training data used · 1 source · 1 quote
Video frame extraction at 1 FPS used · 1 source · 1 quote
Web knowledge graph construction and question generation · web knowledge graph construction used · 1 source · 1 quote
Video sampling parameters · mm_processor_kwargs optional · 1 source · 1 quote

data filtering 37

Conservative model-based noise filtering · model-based and deliberately conservative core · 1 source · 1 quote
Continuous contribution-score ranking core · 1 source · 1 quote
Multi-dimensional document quality scoring core · 1 source · 1 quote
Quality-bucket sampling · samples from corresponding quality buckets core · 1 source · 1 quote
CSAM filtering · child sexual abuse material filtering used · 3 sources · 3 quotes
Sensitive data filtering used · 3 sources · 3 quotes
CBRN pre-training data filtering · CBRN safety filtering of pre-training data, CBRN safety pre-training data filtering used · 2 sources · 2 quotes
Filtering batched auto-generated and templated content · Quality filtering against model collapse used · 2 sources · 2 quotes
Heuristic filtering · Heuristic filtering for multilingual data used · 2 sources · 2 quotes
Keyword- and regex-based filtering · Keyword- and regex-based content filtering, targeted keyword- and regex-based filters used · 2 sources · 2 quotes
Model-based refusal and inability filtering used · 2 sources · 2 quotes
Pass@100 filtering used · 2 sources · 2 quotes
Pathological repetition filtering · n-gram repetition filtering used · 2 sources · 2 quotes
Unified data filtering pipeline used · 2 sources · 2 quotes
Automated verification used · 1 source · 1 quote
Content quality and safety filtering used · 1 source · 1 quote
Data cleaning · fine-grained data cleaning used · 1 source · 1 quote
Data filtering pipeline used · 1 source · 1 quote
Dense annotation of ambiguous low-quality data · densely annotate this region used · 1 source · 1 quote
Image-text relevance filtering · image-text relevance threshold used · 1 source · 1 quote
Long-context data cleaning pipeline used · 1 source · 1 quote
Multi-dimensional data filtering used · 1 source · 1 quote
PCA-based dimension selection for decorrelated quality signals · PCA-informed subset of its dimensions used · 1 source · 1 quote
Permissive-license filtering used · 1 source · 1 quote
Post-training data filtering for safety and factuality · post-training data filtering used · 1 source · 1 quote
Pre-training data filtering for decontamination and safety · data filtering used · 1 source · 1 quote
Propella used · 1 source · 1 quote
SmolVLM-based image-text quality scoring · SmolVLM used · 1 source · 2 quotes
Structural checks for malformed examples · structural checks used · 1 source · 1 quote
Three-stage question filtering used · 1 source · 1 quote
Two-axis document-quality labeling by noise and information · two complementary axes used · 1 source · 1 quote
One-task-per-repository diversity filtering optional · 1 source · 1 quote

deduplication 5

Deduplication used · 2 sources · 2 quotes
In-pack deduplication · in-pack deduplication constraint used · 1 source · 1 quote
Semantic image deduplication · deduplicate them based on image semantics used · 1 source · 1 quote
Snapshot-level fuzzy deduplication used · 1 source · 1 quote

synthetic data 56

Agentic data synthesis pipeline core · 4 sources · 4 quotes
Automatic synthesis of task-oriented RL environments · automatic environment-synthesis agent, Large-Scale Agentic Tasks core · 2 sources · 2 quotes
Modular synthetic-data pipeline composition · modular approach to synthesis core · 1 source · 1 quote
Knowledge distillation for synthetic data · knowledge distillation, distillation used · 3 sources · 3 quotes
Corpus rephrasing with fidelity verification · rephrasing recipe used · 1 source · 1 quote
End-to-end synthetic environment generation · Synthetic environment generation pipelines used · 1 source · 1 quote
Heterogeneous answer-generation agents · answer-generation agents used · 1 source · 1 quote
Iterative task-difficulty escalation used · 1 source · 1 quote
Knowledge-graph-guided task synthesis used · 1 source · 1 quote
Metadata-conditioned synthetic generation · Help LLMs through metadata used · 1 source · 1 quote
Mocked tools for agent environments used · 1 source · 1 quote
Multi-agent task-attempt validation · multiple distinct agents attempt the task used · 1 source · 1 quote
Multi-stage synthetic-data generation cascade · Multi-stage cascade used · 1 source · 1 quote
Programmatic multimodal data generation · programmatic multimodal data used · 1 source · 1 quote
Reward-driven task-construction training used · 1 source · 1 quote
Rubric-guided RL data used · 1 source · 1 quote
Search-based question construction · question-construction agent used · 1 source · 1 quote
Search-capable answer verification agent · verification agent used · 1 source · 1 quote
Synthetic data generation used · 1 source · 1 quote
Synthetic data generation and augmentation · synthetically generated or augmented used · 1 source · 1 quote
Synthetic privacy-preserving personas used · 1 source · 1 quote
Synthetic system-message augmentation · synthetically generated system messages used · 1 source · 1 quote
Targeted synthetic data rephrasing · targeted synthetic rephrasing at scale used · 1 source · 1 quote
Task-specific data injection during mid-training · Task-specific data injection used · 1 source · 1 quote
Tool-set and schema randomization used · 1 source · 1 quote
Two-sided test validation for synthetic coding tasks · two-sided correctness check used · 1 source · 1 quote
Counterfactual data augmentation mentioned · 2 sources · 2 quotes

data mixture & curriculum 26

AutoMixer · Automated data-mixture optimization core · 1 source · 2 quotes
Sample Mixer core · 1 source · 1 quote
Vulnerability-discovery data inclusion · added vulnerability-discovery data used · 2 sources · 2 quotes
Agent-centric multimodal data mixture · agent-centric data mixture used · 1 source · 1 quote
Blind pairwise bucket-boundary calibration used · 1 source · 1 quote
Data Scheduler used · 1 source · 1 quote
Deterministic pre-splitting of ultra-long documents · deterministically pre-split used · 1 source · 2 quotes
Domain-specific multimodal datasets · domain-specific datasets used · 1 source · 1 quote
Dynamic sampling for mixed-task reinforcement learning · Dynamic sampling for mixed-task RL, dynamic sampler used · 1 source · 1 quote
Gaussian-based data mixture and curriculum construction · Gaussian-based approach used · 1 source · 1 quote
Interleaved data used · 1 source · 1 quote
Interleaved multimodal training · interleaved training used · 1 source · 1 quote
KL-regularized data-mixture optimization used · 1 source · 1 quote
Long-context data upsampling · upsampling long-context data used · 1 source · 1 quote
Pass-rate-based RL task sampling · pass-rate buckets used · 1 source · 1 quote
Smooth weighted round-robin source scheduling · smooth weighted round-robin used · 1 source · 1 quote
Text-only pre-training corpus used · 1 source · 1 quote
Three-stage data mixture strategy used · 1 source · 1 quote
Three-stage pre-training curriculum · three distinct stages used · 1 source · 1 quote
Two-phase curriculum used · 1 source · 1 quote
Union-based integration of text-only and multimodal corpora · union of both data sources used · 1 source · 1 quote
ProtocolQA-aligned RL training datasets · Aligned RL datasets with ProtocolQA evaluated · 1 source · 1 quote

sequence packing 6

Length-aware best-fit packing · length-aware best-fit packing strategy used · 2 sources · 2 quotes
Best-fit packing · best-fit packing algorithm, best-fit sequence packing used · 1 source · 2 quotes
Cross-source document packing used · 1 source · 1 quote
Group-local packing used · 1 source · 1 quote
Interleaved image-text sequences · interleaved image-text data construction used · 1 source · 1 quote
Sequence packing used · 1 source · 1 quote

tokenization 8

XTML (eXtensible Token Markup Language) core · 1 source · 1 quote
o200k_harmony tokenizer · o200k_harmony Byte Pair Encoding tokenizer used · 2 sources · 2 quotes
Quick Instruction tokens · Quick Instruction used · 2 sources · 2 quotes
Byte-level Byte Pair Encoding (BPE) used · 1 source · 1 quote
SentencePiece tokenizer used · 1 source · 1 quote
Separate PT and IT end tokens · PT versus IT formatting used · 1 source · 1 quote
Subtoken averaging used · 1 source · 1 quote

post-training 267

filed at the root 10

Scalable RL at Agent Scale core · 3 sources · 3 quotes
Scalable reinforcement learning post-training · Scalable Reinforcement Learning Framework core · 2 sources · 2 quotes
Reinforcement learning for low-pass-rate tasks · RL reserved for tasks with low pass rate core · 1 source · 2 quotes
Three-stage post-training with multi-teacher on-policy distillation · three-stage post-training paradigm core · 1 source · 1 quote
SFT followed by RL and on-policy distillation default · 1 source · 2 quotes
Extended post-training used · 1 source · 1 quote
Joint SFT and RL used · 1 source · 1 quote
Post-training on diverse domains used · 1 source · 1 quote
Post-training optimization used · 1 source · 1 quote
Three-stage Thinker post-training strategy · three-stage strategy for the Thinker used · 1 source · 1 quote

supervised fine-tuning 24

LoRA fine-tuning · LoRA adapter fine-tuning, LoRA adapter core · 2 sources · 2 quotes
Light SFT on teacher-distribution data · teacher-distribution SFT warmup core · 1 source · 2 quotes
SFT checkpoint for RL research · SFT starting point for RL research core · 1 source · 1 quote
Supervised fine-tuning · SFT, Supervised Fine Tuning (SFT) used · 14 sources · 17 quotes
Rejection sampling · rejection sampling fine-tuning used · 2 sources · 2 quotes
Self-correction cold start · Cold start from self-correction used · 2 sources · 2 quotes
SFT bootstrapping with synthetic data used · 2 sources · 2 quotes
Adversarial fine-tuning used · 1 source · 1 quote
Autoregressive fine-tuning used · 1 source · 1 quote
Cold-started SFT model used · 1 source · 1 quote
Evaluation-based early stopping · evaluation-based early stopping during SFT, early stopping based on evaluation scores used · 1 source · 1 quote
Instruction hierarchy training used · 1 source · 1 quote
Instruction-following fine-tuning used · 1 source · 1 quote
Lightweight speaker fine-tuning used · 1 source · 1 quote
Prompt-based cold start for tool use · Cold-Start used · 1 source · 1 quote
Reasoning-data system prompt used · 1 source · 1 quote
Safety training to an internal specification · Safety training to internal spec used · 1 source · 1 quote
SFT on MiMo-generated task data used · 1 source · 1 quote
Specialist-model training · Specialist training, independent specialist models used · 1 source · 1 quote
Token-budget-based SFT data blending · token-level target proportions used · 1 source · 1 quote
Fine-tuning · fine-tune optional · 3 sources · 3 quotes

reinforcement learning algorithm 62

Group Relative Policy Optimization · GRPO, GRPO (Group Relative Policy Optimization) core · 15 sources · 17 quotes
Groupwise Advantage Redistribution · Groupwise Advantage Redistribution (GAR) core · 3 sources · 4 quotes
Chain-of-Thought Reinforcement Learning · similar CoT RL techniques as OpenAI o3 core · 2 sources · 2 quotes
Mixed Reinforcement Learning · Mixed RL Training core · 2 sources · 2 quotes
Effort-Conditioned Reinforcement Learning core · 1 source · 1 quote
Keep Routing core · 1 source · 1 quote
Penalty Module core · 1 source · 1 quote
Dropping All-Zero-Advantage Groups default · 1 source · 1 quote
Reinforcement Learning · RL, Reinforcement Learning (RL) used · 11 sources · 12 quotes
IcePop · IcePop technique, token-level masking using IcePop strategy used · 2 sources · 3 quotes
Keep Sampling Mask used · 2 sources · 2 quotes
Off-Policy Sequence Masking · Off-policy sequence masking for GRPO used · 2 sources · 3 quotes
Reinforcement Learning from Verifiable Rewards · unified RLVR, RLVR used · 2 sources · 3 quotes
Reinforcement Learning Post-Training · RL post-training, larger-scale RL post-training used · 2 sources · 2 quotes
Rollout Routing Replay · Rollout Routing Replay (R3) used · 2 sources · 3 quotes
Unbiased KL Estimate used · 2 sources · 3 quotes
Abstention Training · abstention-focused reinforcement learning used · 1 source · 1 quote
Advantage Shaping · Advantage shaping on flagged tokens used · 1 source · 1 quote
Agentic RL Task Mix used · 1 source · 1 quote
Concatenated Routing Replay used · 1 source · 1 quote
Derived-Latency Penalty used · 1 source · 1 quote
Direct Double-Sided Importance Sampling used · 1 source · 2 quotes
Domain-Specialized RL Experts · scaling RL across three broad domains used · 1 source · 1 quote
Domain-Specific GRPO Training used · 1 source · 1 quote
Dynamic Abstention-Reward Calibration used · 1 source · 1 quote
Dynamic Sampling used · 1 source · 1 quote
Effort-Dependent Length Penalty used · 1 source · 1 quote
Flagged-Token Masking and Penalty · flagged tokens used · 1 source · 1 quote
Freezing the MoE Router During Reinforcement Learning · freeze the router for RL training, freeze the MoE router used · 1 source · 3 quotes
Group-Relative Length Penalty used · 1 source · 1 quote
Group-Wise Policy Optimization · group-wise policy optimization algorithm used · 1 source · 1 quote
GRPO with Masked Importance Sampling used · 1 source · 1 quote
GSPO used · 1 source · 1 quote
Hierarchical Penalty Escalation · Penalties escalate along the hierarchy used · 1 source · 1 quote
Incremental Reinforcement Learning used · 1 source · 1 quote
Interaction-Aligned Reinforcement Learning · Interaction-Aligned RL used · 1 source · 1 quote
Keep Candidate-Set Replay · candidate-set replay used · 1 source · 1 quote
Large-Scale Reinforcement Learning · Reinforcement learning for post-training used · 1 source · 1 quote
Per-Token Tool-Error Step Penalty · Tool-error step penalty used · 1 source · 1 quote
Pivot Reinforcement Learning · PivotRL used · 1 source · 2 quotes
PPO-Style Policy-Ratio Clipping · PPO-style clipping used · 1 source · 1 quote
Progressively Complex Task Distributions used · 1 source · 1 quote
Quality-Weighted Advantage Redistribution · quality factors used · 1 source · 1 quote
Reinforcement Learning from Code Execution Feedback · RLCEF tasks used · 1 source · 1 quote
Reinforcement Learning Training · RL training used · 1 source · 1 quote
RL with Rubric and Claims Graders used · 1 source · 1 quote
SAO with Compaction · SAO with compaction (RL strategy) used · 1 source · 1 quote
Single Mixed Reinforcement-Learning Run used · 1 source · 1 quote
Single-Harness Reinforcement Learning · single-harness RL used · 1 source · 1 quote
Moonlight Scaling · Moonlight scaling in RL optional · 1 source · 1 quote

reward modelling 35

Groupwise Agentic Grading core · 3 sources · 4 quotes
Groupwise Reward Synthesis · Groupwise Reward Synthesis (GRS) core · 3 sources · 4 quotes
Binary task verifier core · 1 source · 1 quote
Binary terminal-verifier reward · binary verifier core · 1 source · 1 quote
Deterministic chain of checkers core · 1 source · 1 quote
Generative Reward Model · GenRM, Generative Reward Model (GRM) used · 6 sources · 7 quotes
Adversarial screening used · 2 sources · 2 quotes
Length-adjusted RL reward · length penalty reward, length penalty used · 2 sources · 2 quotes
Verifier cross-checking · verifier cross-checks used · 2 sources · 2 quotes
Abstention-aware reward for factual QA · Abstention-aware rewards for factual QA used · 1 source · 1 quote
Agentic Generative Reward Model · Agentic Generative Reward Model (GRM) used · 1 source · 1 quote
Behavior rubrics used · 1 source · 1 quote
Collaboration bonus · Collaboration bonus in RL reward used · 1 source · 1 quote
Hack-agent screening · Hack Agent used · 1 source · 1 quote
Hybrid reward system used · 1 source · 1 quote
Language consistency reward used · 1 source · 1 quote
Monitoring-only penalty strategy · monitor used · 1 source · 1 quote
Multi-level reward formulation used · 1 source · 1 quote
Multiplicative reward synthesis used · 1 source · 1 quote
Outcome Reward Model · Outcome Reward Models, Outcome Reward Models (ORMs) used · 1 source · 1 quote
Per-token tool-error reward shaping used · 1 source · 1 quote
Principle-conditioned Generative Reward Model · principle-following GenRM used · 1 source · 1 quote
Public and hidden verifier pairing used · 1 source · 1 quote
Rule-based or model-judged trajectory error detection · Rule used · 1 source · 1 quote
Rule-based outcome reward · rule-based outcome rewards used · 1 source · 1 quote
Rule-based verifier · rule-based verifiers used · 1 source · 1 quote
Segment-level behavioral penalties used · 1 source · 1 quote
Solution rubrics used · 1 source · 1 quote
Test-difficulty-driven code reward used · 1 source · 1 quote
Training-time trajectory auditing · Training-Time Auditing used · 1 source · 1 quote

preference optimization 4

Reinforcement Learning from Human Feedback · RLHF used · 3 sources · 3 quotes
Budget-based verbosity control used · 1 source · 1 quote
Deliberative alignment used · 1 source · 1 quote
Direct Preference Optimization · Direct Preference Optimization (DPO) used · 1 source · 1 quote

policy distillation 19

Multi-Teacher On-Policy Distillation · Mixture of On-Policy Distillation, MOPD core · 12 sources · 17 quotes
On-Policy Distillation · on-policy distillation (OPD) core · 7 sources · 9 quotes
Specialist Distillation · multi-domain specialist distillation, specialist-model distillation used · 3 sources · 3 quotes
Prefix-Conditioned On-Policy Distillation · Prefix-Conditioned OPD, SFT-Prefix OPD used · 2 sources · 5 quotes
Asynchronous Multi-Teacher On-Policy Distillation · asynchronous MOPD used · 1 source · 1 quote
Autonomous Student Rollouts · Standard MOPD used · 1 source · 2 quotes
Behavior–Proximal Policy Decoupling used · 1 source · 1 quote
Distillation Fine-Tuning on MiMo-Generated Data · trained on MiMo-generated data used · 1 source · 1 quote
IcePop Token-Level Loss Masking · token-level masking using IcePop strategy used · 1 source · 1 quote
Large-Scale On-Policy Distillation used · 1 source · 2 quotes
Model Distillation · distillation used · 1 source · 1 quote
Multi-Objective Policy Distillation · MOPD used · 1 source · 1 quote
On-Policy Cross-Stage Distillation used · 1 source · 2 quotes
Per-Token On-Policy Distillation Reward · per-token OPD reward used · 1 source · 1 quote
SFT–RL–On-Policy Distillation Pipeline used · 1 source · 1 quote
Teacher-Trajectory and SFT-History Reuse used · 1 source · 1 quote
Off-Policy Distillation not used · 1 source · 1 quote

rollout & RL infrastructure 76

Asynchronous reinforcement learning · asynchronous RL architecture, asynchronous training paradigm core · 6 sources · 7 quotes
Agent Loop core · 1 source · 1 quote
Agent-centric rollout execution · agent-centric execution model core · 1 source · 1 quote
Asynchronous reinforcement learning framework · asynchronous RL framework core · 1 source · 1 quote
Harness Pool core · 1 source · 1 quote
Hierarchical trajectory data organization core · 1 source · 1 quote
Large-scale asynchronous RL core · 1 source · 1 quote
One-step off-policy asynchronous reinforcement learning · one-step off-policy asynchronous RL setup core · 1 source · 2 quotes
Payload Porter core · 1 source · 1 quote
Predictive Rollout Dispatch core · 1 source · 2 quotes
Sample-level dispatch default · 1 source · 1 quote
Partial rollout · Partial rollout for mixed-task RL, partial rollout for synchronous RL used · 3 sources · 5 quotes
Asynchronous reinforcement learning infrastructure · asynchronous RL infrastructure used · 2 sources · 2 quotes
Co-located RL training used · 2 sources · 2 quotes
Seamless Rollout Engine used · 2 sources · 2 quotes
SLIME · SLIME (Reinforcement Learning framework) used · 2 sources · 2 quotes
Token-in-token-out (TITO) · token-in-token-out, token-in, token-out (TITO) API design used · 2 sources · 3 quotes
Adaptive Rollout Concurrency used · 1 source · 1 quote
Adaptive Rollout Scheduling used · 1 source · 1 quote
Asynchronous Agent RL algorithms used · 1 source · 1 quote
Asynchronous sample generation · asynchronous generation of samples used · 1 source · 1 quote
Blocking in-flight rollout steps on weight updates · blocking in-flight rollout steps on update used · 1 source · 1 quote
Bounding the maximum off-policy ratio · bound the maximum off-policy ratio used · 1 source · 1 quote
Capping trajectory staleness used · 1 source · 1 quote
Closed-loop multi-turn rollout · Multi-turn rollout used · 1 source · 1 quote
Co-located long-context agentic RL system used · 1 source · 1 quote
Composite early-stop strategy · early stop strategy used · 1 source · 1 quote
Compute-node sharding into scale units used · 1 source · 1 quote
Consistent configuration transitions used · 1 source · 1 quote
Data re-sampling strategy used · 1 source · 1 quote
Data Scheduler used · 1 source · 2 quotes
Decoupled control plane and data plane · Disaggregated Data Plane and Control Plane used · 1 source · 2 quotes
Dialogue prefix matching · Prefix matching used · 1 source · 1 quote
Discarding early-returned short samples used · 1 source · 1 quote
Dynamic training recipe reconfiguration · dynamic reconfiguration during training used · 1 source · 1 quote
Full-lifecycle monitoring of RL tasks used · 1 source · 1 quote
Heterogeneous Agent Harnesses used · 1 source · 1 quote
Incremental image transfer used · 1 source · 1 quote
Incremental multimodal-delta transfer used · 1 source · 1 quote
Isolated resettable rollout sandboxes used · 1 source · 1 quote
Long-horizon rollout budgets · long-horizon rollout budgets in RL used · 1 source · 1 quote
Mixed-task rollout infrastructure used · 1 source · 1 quote
Multi-harness rollouts · multi-harness rollout in RL used · 1 source · 1 quote
Multi-Task Rollout Orchestrator used · 1 source · 1 quote
Multimodal rollout data handling · Multi-Modal Data used · 1 source · 1 quote
On-policy rollouts used · 1 source · 1 quote
Per-dataset rollout concurrency limiting used · 1 source · 1 quote
Per-source oversampling allocation · oversampling ratio used · 1 source · 1 quote
Per-source priors for rollout sequence-length estimation · per-source priors used · 1 source · 1 quote
Persistent host-actor pools · fixed-size pools of persistent host actors used · 1 source · 1 quote
Preemptible rollout service used · 1 source · 1 quote
Sample Mixer · A sample mixing mechanism used · 1 source · 1 quote
Sample replay used · 1 source · 1 quote
Sample-grained garbage collection used · 1 source · 2 quotes
Scaffold-agnostic rollout control layer used · 1 source · 1 quote
Token-level interruption · Token-level interruption of generation used · 1 source · 2 quotes
Tool Manager used · 1 source · 2 quotes
Toolbox used · 1 source · 2 quotes
Asynchronous group-wise grading optional · 1 source · 1 quote
Deficit-corrected scheduling evaluated · 1 source · 2 quotes
Batch-level rollout dispatch · batch-level dispatch of rollout prompts not used · 1 source · 1 quote
Prompt-level dispatch · prompt-level dispatch of rollout prompts not used · 1 source · 1 quote

agentic post-training 31

Agentic reinforcement learning · agentic RL, Large-scale agentic reinforcement learning core · 4 sources · 4 quotes
Multi-environment reinforcement learning core · 1 source · 1 quote
Reinforcement learning in an agent harness core · 1 source · 1 quote
Agentic tool-use training · agentic tool-use post-training used · 2 sources · 2 quotes
Environment hardening used · 2 sources · 2 quotes
Agentic post-training used · 1 source · 1 quote
Autonomous Execution Tasks (AET) used · 1 source · 1 quote
Concurrent multi-environment post-training used · 1 source · 1 quote
Container-level network isolation used · 1 source · 1 quote
Environment preparation to prevent solution leakage · Environment Preparation used · 1 source · 1 quote
Generate-verify-refine loop used · 1 source · 1 quote
Harness-optimized training used · 1 source · 1 quote
Multi-environment reinforcement learning from verifiable rewards · multi-environment RLVR used · 1 source · 1 quote
Multi-harness training · Multi-harness reinforcement learning used · 1 source · 3 quotes
Multi-turn collaboration simulator used · 1 source · 1 quote
Python tool use in chain-of-thought · python tool used · 1 source · 1 quote
Re-post-training for agentic capabilities used · 1 source · 1 quote
Reinforcement learning on synthetic agentic data · large-scale RL on synthetic data used · 1 source · 2 quotes
Repair-agent environment correction loop used · 1 source · 2 quotes
Seed-based terminal task generation used · 1 source · 1 quote
Training agentic policies with real-world tools · real-world tools used · 1 source · 1 quote
Unified trajectory representation for agentic reinforcement learning · unified trajectory representation used · 1 source · 1 quote
Unified white-box reinforcement-learning environment · unified white-box RL environment used · 1 source · 1 quote
Reinforcement learning restricted to search and code environments · RL only in search and code environments evaluated · 1 source · 1 quote

mid-training & continual pretraining 6

Continued training for visual understanding · continued training core · 1 source · 1 quote
Continual pretraining · continued pre-training, Continuous Pretraining used · 3 sources · 3 quotes
Continual pretraining for long-context extension · long-context extension, long-context training used · 3 sources · 3 quotes
Agent-centric mid-training used · 1 source · 1 quote
Mid-training alignment data used · 1 source · 1 quote
Mid-training phase used · 1 source · 1 quote

inference & serving 375

filed at the root 1

communication optimization used · 1 source · 1 quote

decoding strategy 28

Speculative decoding · speculative decoding methods, speculative decoding module core · 22 sources · 25 quotes
Multi-Token Prediction · MTP-based speculative decoding, Multi-Token Prediction (MTP) core · 20 sources · 20 quotes
DSpark · DSpark speculative decoding, DSpark speculative-decoding module core · 8 sources · 10 quotes
DFlash · Block-6 DFlash speculative decoding, block-6 DFlash for speculative decoding core · 6 sources · 8 quotes
Fused recurrent replay kernel · single fused kernel core · 1 source · 1 quote
Recursive shared MTP-head drafting core · 1 source · 2 quotes
Standardized sampling configuration default · 3 sources · 3 quotes
Same-checkpoint target and draft weights default · 1 source · 1 quote
EAGLE · EAGLE (speculative decoding), EAGLE (speculative decoding algorithm) used · 9 sources · 9 quotes
NEXTN speculative decoding · native NEXTN speculative decoding, NEXTN speculative decoding algorithm used · 4 sources · 4 quotes
Presence Penalty · Presence penalty adjustment, Presence-penalty tuning used · 4 sources · 5 quotes
Best-of-N scaffolding · Best of K scaffolding, Best-of-n selection used · 3 sources · 4 quotes
Multi-stage candidate filtering · multi-stage filtering pipeline used · 2 sources · 2 quotes
EAGLE-3-style draft-model fine-tuning · Draft Model Fine-Tuning used · 1 source · 1 quote
Grammar-constrained decoding used · 1 source · 1 quote
Longest-trace selection · longest thinking trace selection used · 1 source · 1 quote
MTP-1 speculative decoding used · 1 source · 1 quote
Throughput-based draft-block sizing used · 1 source · 1 quote
Top-k over token clusters · top-k operation on clusters of tokens used · 1 source · 1 quote
Top-p and top-k sampling used · 1 source · 1 quote
Multi-layer EAGLE · enable-multi-layer-eagle optional · 3 sources · 3 quotes
Speculative sampling optional · 3 sources · 3 quotes
Task-specific sampling parameters · sampling parameters optional · 2 sources · 2 quotes
Chat Prefix Completion optional · 1 source · 1 quote
Concurrency-aware draft-length tuning evaluated · 1 source · 1 quote

reasoning control 42

Configurable reasoning effort · dial 'thinking effort' up or down, reasoning_effort request field core · 33 sources · 34 quotes
Chain-of-thought reasoning · Chain-of-Thought, COT core · 8 sources · 8 quotes
Default thinking mode · Default thinking mode for generation, thinking mode by default core · 8 sources · 8 quotes
Interleaved thinking between tool calls · Interleaved thinking, per-request control via enable_thinking core · 3 sources · 4 quotes
Always-on thinking mode · always-on thinking block, always has thinking enabled core · 2 sources · 2 quotes
Control-token-enabled thinking mode · configurable thinking modes core · 2 sources · 2 quotes
Capped linear reasoning-token length deduction · length deduction core · 1 source · 1 quote
Effort-dependent exponential token-penalty schedule · effort-dependent token-penalty coefficient, exponential token penalty schedule core · 1 source · 2 quotes
Thinking mode selection · Thinking mode, thinking default · 4 sources · 4 quotes
Maximum thinking effort · Max thinking effort, max thinking default · 3 sources · 3 quotes
clear_thinking chat-template parameter · clear_thinking chat template flag, clear_thinking default · 2 sources · 2 quotes
Deployment-time scalar effort control · Scalar effort control of response length, scalar efforts default · 1 source · 5 quotes
Cross-turn persistent reasoning history · persistent reasoning-history chat template, persistent thinking history used · 3 sources · 3 quotes
Inference-time reasoning budget control · inference-time budget control used · 3 sources · 4 quotes
Reasoning parser · reasoning parser step3p5 used · 3 sources · 3 quotes
Generate-verify-refine loop · Generate-verify-refine test-time scaling, generate–verify–refine methodology used · 2 sources · 2 quotes
Medium-effort reasoning mode used · 2 sources · 2 quotes
Test-time compute scaling · test-time scaling, test-time-scaling framework used · 2 sources · 2 quotes
Per-problem reasoning-budget control · per-problem budget control mechanism used · 1 source · 1 quote
Think-tag response formatting used · 1 source · 1 quote
Thinking mode selection via generation prefix · channels used · 1 source · 1 quote
Thinking modes (off and max) · two modes: off and max used · 1 source · 1 quote
Turn budget capping · turn budgets used · 1 source · 1 quote
Turn-aware prompting used · 1 source · 1 quote
Disabling reasoning via chat-template configuration · disabling reasoning mode via chat template, Disable Reasoning optional · 5 sources · 5 quotes
Adaptive reasoning optional · 1 source · 1 quote
Configurable thinking or reasoning mode · configurable thinking/reasoning mode optional · 1 source · 1 quote
reasoning-enabled inference toggle · reasoning enabled boolean optional · 1 source · 1 quote
Step-by-step thinking mode optional · 1 source · 1 quote
Task-appropriate maximum output length · Adequate Output Length optional · 1 source · 1 quote
Parallel-fewest-step sampling · parallel scaling, Parallel-fewest-step evaluated · 2 sources · 2 quotes
Parallel test-time compute scaling · in parallel evaluated · 1 source · 1 quote
Reward-optimized preferred reasoning length · preferred reasoning length evaluated · 1 source · 1 quote
Serial test-time compute scaling through context management · serially through context management evaluated · 1 source · 1 quote
Qwen3 soft thinking switch not used · 2 sources · 2 quotes
User-configurable thinking-effort control not used · 1 source · 1 quote

KV cache management 42

Compressed KV caching · KV cache compression, Smaller KV cache core · 2 sources · 2 quotes
SWA Bounded Replay core · 2 sources · 4 quotes
Block-based KV cache · block-based key-value cache core · 1 source · 1 quote
Cross-layer KV-cache reuse · Cross-layer key-value cache reuse core · 1 source · 1 quote
Customized heterogeneous KV-cache layout · customized KV cache layout core · 1 source · 1 quote
Fine-grained prefix hashing · Prefix hashing runs on fine hash blocks core · 1 source · 1 quote
KDA-aware prefix-cache management · joint KDA–MLA prefix cache management, KDA-aware prefix cache core · 1 source · 2 quotes
Persistent per-dialogue-context KV caching · Context Caching core · 1 source · 1 quote
Projected-input caching for speculative KDA rollback · cache only these projected inputs core · 1 source · 1 quote
Shared-free-list cache allocation core · 1 source · 1 quote
Write-back external KV-cache policy · write-back design core · 1 source · 1 quote
Chunked prefill · --chunked-prefill-size 16384, chunked prefilling default · 6 sources · 6 quotes
Automatic cache · Full automatic Cache support default · 1 source · 1 quote
Prefix caching · prefix cache used · 5 sources · 5 quotes
Prompt caching used · 4 sources · 4 quotes
Language-model-only serving mode · language-model-only, Text-Only used · 3 sources · 3 quotes
RadixCache · radix cache used · 2 sources · 2 quotes
Disable prefix caching for benchmarking · disabling prefix caching used · 1 source · 1 quote
Encoder SWA bounded replay used · 1 source · 1 quote
Hierarchical KV caching · hierarchical extension of the cache used · 1 source · 1 quote
Inference-side KV-cache reset on weight synchronization · KV-cache reset on weight sync used · 1 source · 1 quote
KDA with prefill cache used · 1 source · 1 quote
Key-value reuse in global attention layers · reuse of keys as values in global layers used · 1 source · 1 quote
KV-cache sharing used · 1 source · 1 quote
Least-recently-used eviction · LRU eviction policy used · 1 source · 1 quote
On-disk KV-cache storage · on-disk KV cache storage mechanism used · 1 source · 1 quote
Overlapped host-memory prefetching · host-memory prefetching used · 1 source · 1 quote
Persistent KV-cache management used · 1 source · 1 quote
Pre-scheduled tile chunking used · 1 source · 1 quote
Request-level prefix cache used · 1 source · 1 quote
Stateless in-memory prefix caching · stateless prefix caching in memory used · 1 source · 1 quote
Periodic cache checkpointing · cache checkpointing, Periodic checkpointing of SWA KV entries optional · 3 sources · 3 quotes
Zero SWA caching optional · 2 sources · 2 quotes
8-bit Mamba cache quantization evaluated · 1 source · 1 quote
Mamba prefix caching in align mode · prefix caching in align mode for Mamba evaluated · 1 source · 1 quote

inference quantization 64

Quantization · quantization algorithms, quantized versions core · 3 sources · 3 quotes
FP4 KV-cache quantization · FP4 main KV cache, FP4 KV caching core · 2 sources · 5 quotes
Dynamic activation scaling core · 1 source · 1 quote
Mixed-FP8 layers in an NVFP4 recipe · mixed-FP8 layers core · 1 source · 1 quote
Mobile-specialized quantization schema · custom mobile-quantization schema core · 1 source · 1 quote
NVFP4 quantization for routed-expert GEMMs · NVFP4 routed-expert GEMMs core · 1 source · 1 quote
Offline weight-layout permutation · weight layout is permuted offline core · 1 source · 1 quote
Per-tensor FP8 quantization core · 1 source · 1 quote
Selective retention of BF16 precision core · 1 source · 1 quote
FP8 · FP8 (8-bit floating point), FP8 quantization default · 14 sources · 14 quotes
NVFP4 quantization · NVFP4 4-bit floating-point quantization, NVFP4 checkpoint default · 5 sources · 5 quotes
NVFP4 KV-cache quantization · NVFP4 KV cache, NVFP4 KV caching default · 1 source · 2 quotes
Post-RoPE KV-cache quantization · quantizing the cache after RoPE default · 1 source · 1 quote
FP8 KV-cache quantization · FP8 KV cache during RL rollouts, FP8 E4M3 key-value cache used · 13 sources · 15 quotes
MXFP4 weight quantization · Microscaling FP4 (MXFP4) weights, MXFP4 weights used · 3 sources · 3 quotes
Post-training quantization · Post-Training Quantization (PTQ) used · 3 sources · 3 quotes
Four-Over-Six · Four-Over-Six FP4 weight scale selection, Four-Over-Six scale selection method used · 2 sources · 3 quotes
GGUF · GGUF (GPT-Generated Unified Format), GGUF quantization used · 2 sources · 2 quotes
MXFP8 activation quantization · Microscaling FP8 (MXFP8) activations, MXFP8 activations used · 2 sources · 2 quotes
SSM cache quantization · quantize the Mamba SSM cache used · 2 sources · 2 quotes
AWQ INT4 weight quantization (W4A16) · INT4 AWQ weight quantization used · 1 source · 1 quote
BF16 inference used · 1 source · 1 quote
Block-wise E4M3 FP8 weight quantization · FP8 quantization with block-wise e4m3, native FP8 (block-wise e4m3) weights used · 1 source · 1 quote
Channel-wise quantization used · 1 source · 1 quote
Embedding and KV-cache quantization · Embedding and KV cache optimization used · 1 source · 1 quote
Flex_AWQ_SSZ used · 1 source · 1 quote
FP16 cache storage with stochastic rounding · FP16 Cache with Stochastic Rounding used · 1 source · 1 quote
FP16 multimodal projector used · 1 source · 1 quote
FP4 precision for the attention indexer · FP4 precision for indexer, FP4 precision for attention indexer used · 1 source · 1 quote
FP8 inference · FP8 rollouts used · 1 source · 1 quote
FP8 low-precision speculative draft computation · Low-precision draft computation (FP8) used · 1 source · 1 quote
Heuristic mixed per-layer precision quantization · heuristic mixed per-layer precision recipe used · 1 source · 1 quote
Mixed-precision INT4/INT8 layer-wise quantization · mixed-precision quantization strategy used · 1 source · 1 quote
ModelOpt FP4 quantization used · 1 source · 1 quote
MXFP4 tensor packing · pack every two values in one uint8 value used · 1 source · 1 quote
MXFP8 block-scale regrouping at load used · 1 source · 1 quote
NVFP4 quantization with modelopt · NVFP4 model used · 1 source · 1 quote
NVIDIA ModelOpt quantization · modelopt quantization used · 1 source · 1 quote
Post-quantization of MoE model layers · quantization of MoE layers used · 1 source · 1 quote
QuaRot used · 1 source · 1 quote
Random Hadamard transform · Random Hadamard Transforms used · 1 source · 1 quote
SpinQuant R1 rotation used · 1 source · 1 quote
Static activation quantization · Static activations used · 1 source · 1 quote
Targeted 2-bit quantization used · 1 source · 1 quote
W4A16 quantization · W4A16 (weights NVFP4, activations BF16), W4A16 used · 1 source · 1 quote
W4A8 quantization · W4A8 mixed-precision quantization strategy used · 1 source · 1 quote
NVFP4 ModelOpt re-quantization · NVFP4 ModelOpt re-quantization (W4A16), ModelOpt re-quantization optional · 2 sources · 2 quotes
Fine-grained FP8 quantization optional · 1 source · 1 quote
IQ4_XS quantization optional · 1 source · 1 quote
Mobile quantization optional · 1 source · 1 quote
Q3_K_L quantization optional · 1 source · 1 quote
Q4_0 quantization · Q4_0 blockwise quantization optional · 1 source · 1 quote
Q4_K_S quantization optional · 1 source · 1 quote
Block-scaled INT8 quantization with stochastic rounding · Block-scaled INT8 quantization evaluated · 2 sources · 2 quotes
FP8 E4M3 quantization evaluated · 2 sources · 2 quotes
Max-based scaling · Max-based quantization scale calibration, Max-based weight scaling evaluated · 2 sources · 2 quotes
MSE-based scaling · MSE-based weight scaling evaluated · 2 sources · 2 quotes
Empirical bits-per-element budget selection · Bits per Element Selection evaluated · 1 source · 1 quote
MSE calibration evaluated · 1 source · 1 quote
Low-bit quantization · low-bit methods mentioned · 1 source · 1 quote
In-flight block-wise FP8 weight quantization · in-flight block-wise weight quantization not used · 1 source · 1 quote

serving parallelism 17

Zero-copy fused token permutation and unpermutation · zero-copy communication, fused per-mute/unpermute operator core · 1 source · 2 quotes
Tensor parallelism (degree 4) · Four-way tensor parallelism for serving, tensor-parallel-size 4 default · 2 sources · 2 quotes
DeepEP · --moe-a2a-backend deepep, DeepEP expert parallelism across nodes used · 4 sources · 4 quotes
Prefill-decode disaggregation · Prefill–Decode (PD) disaggregation used · 4 sources · 4 quotes
Attention data parallelism · Attention Data Parallelism (DP), Data Parallel attention used · 3 sources · 4 quotes
Encoder-Prefill-Decode disaggregation used · 2 sources · 2 quotes
Tensor parallelism (degree 8) · tensor-parallel serving on 8 GPUs, tensor parallel on 8 GPUs used · 2 sources · 2 quotes
Topology-aware NVLink domain placement used · 2 sources · 2 quotes
Data-parallel vision encoding · --mm-encoder-tp-mode data used · 1 source · 1 quote
Tensor parallelism for MoE layers · tensor parallelism in MoE used · 1 source · 1 quote
TensorRT-LLM all-reduce backend · trtllm all-reduce backend used · 1 source · 1 quote
Expert parallelism · enable expert parallelism optional · 1 source · 1 quote
Low-precision MoE combine optional · 1 source · 1 quote
Token migration for balanced expert placement unclear · 1 source · 1 quote

inference scheduling 20

Cross-group pinning of cache-hit blocks core · 1 source · 1 quote
Prefix-cache-aware session affinity scheduling · cache-aware affinity scheduling core · 1 source · 1 quote
Request-class resource-budget admission control · budget-based admission control core · 1 source · 1 quote
Runtime-signal-based rollout concurrency auto-throttling · auto-throttling mechanism core · 1 source · 1 quote
Workload-aware routed-expert GEMM scheduling · workload-aware scheduler core · 1 source · 1 quote
Asynchronous scheduling · async scheduling used · 3 sources · 3 quotes
NUMA binding of workers to GPU-local CPU sockets · NUMA Binding for CPU-GPU affinity, NUMA Binding used · 2 sources · 2 quotes
Envoy-based proxy with custom orchestrator used · 1 source · 1 quote
Greedy feasible-rank placement by remaining capacity · greedy heuristic used · 1 source · 1 quote
Host-side scheduling optimization used · 1 source · 1 quote
Latency-sensitive execution class with priority isolation · a latency-sensitive (LS) execution class used · 1 source · 1 quote
Max sequences tuning · Max sequences tuning for context fitting, --max-num-seqs 32 used · 1 source · 1 quote
Per-node hard admission constraint used · 1 source · 1 quote
Registering NVLink domains as Ray custom resources · registers it as a Ray custom resource used · 1 source · 1 quote
Continuous batching · Continuous batching for inference serving optional · 1 source · 1 quote
Exacto routing · Exacto (highest tool-calling accuracy) optional · 1 source · 1 quote
Increasing batch size for inference · increasing batch size for MoE inference evaluated · 1 source · 1 quote

inference kernel 50

Batch-invariant deterministic kernels · Deterministic and batch-invariant kernels core · 2 sources · 3 quotes
Kernel fusion · inference kernel fusion, operator fusion core · 2 sources · 2 quotes
Fused AttnRes merge and RMSNorm kernel core · 1 source · 1 quote
GPU planning kernel for online expert placement · GPU planning kernel core · 1 source · 1 quote
Separate-stream shared-expert GEMM overlap · separate stream core · 1 source · 1 quote
Shape-aware kernel dispatch for Hyper-Connection · Shape-aware dispatch core · 1 source · 1 quote
Synchronization-free static-shape MoE execution · sync-free MoE execution with static shapes, sync-free execution with static shapes core · 1 source · 2 quotes
WarpDecode token-centric routed-expert decoding kernel · token-centric design of WarpDecode core · 1 source · 1 quote
Sparse pinned-host offload default · 1 source · 1 quote
CUDA Graph · CUDA Graph integration, the step is captured in a CUDA graph used · 3 sources · 3 quotes
CUDA Graph capture size reduction · reduce --max-cudagraph-capture-size used · 2 sources · 2 quotes
FlashAttention 3 · --attention-backend fa3, FlashAttention 3 (fa3) used · 2 sources · 2 quotes
MoE-side chunking used · 2 sources · 2 quotes
TensorRT-LLM multi-head attention backend · trtllm_mha, trtllm_mha attention backend used · 2 sources · 2 quotes
DeepGEMM-based batch-invariant matrix multiplication · replace it end-to-end with DeepGEMM used · 1 source · 1 quote
Dynamic load balancing used · 1 source · 1 quote
Exp-free TopK kernel used · 1 source · 1 quote
Expert-optimized Triton kernels used · 1 source · 1 quote
FA4 sheared-bias attention kernel used · 1 source · 1 quote
FlashAttention used · 1 source · 1 quote
FlashKDA · FlashKDA chunkwise kernel for KDA used · 1 source · 1 quote
Fused mHC kernels · fused kernels of mHC used · 1 source · 1 quote
Fused QSA kernel · Fused sparse-attention and KL-loss kernel used · 1 source · 1 quote
Fused RoPE-attention-RoPE-cast kernel used · 1 source · 1 quote
Hidden-dimension split Combine kernel used · 1 source · 1 quote
Host Codegen used · 1 source · 1 quote
KDA algorithm–system co-design · systems co-design for KDA used · 1 source · 1 quote
Marlin NVFP4 kernels · Marlin NVFP4 linear and MoE kernels used · 1 source · 1 quote
Mega-mHC used · 1 source · 1 quote
Optimized Triton MoE kernel with MXFP4 support · optimized triton MoE kernel used · 1 source · 1 quote
Persistent kernel · persistent kernel rewriting used · 1 source · 1 quote
ROCm AITER sparse MLA attention backend · --attention-backend ROCM_AITER_MLA_SPARSE used · 1 source · 1 quote
Single-Pass mHC used · 1 source · 1 quote
Sparse Flash Attention used · 1 source · 1 quote
Split-K CuTe GEMM for low-batch Hyper-Connection Mix · low-latency split-K CuTe GEMM used · 1 source · 1 quote
Triton MoE kernel used · 1 source · 1 quote
Two-phase forward used · 1 source · 1 quote
FlashAttention 4 · fa4 optional · 1 source · 1 quote
HPC-Ops attention backend · HPC_ATTN optional · 1 source · 1 quote
HPC-Ops fused MoE backend · hpc optional · 1 source · 1 quote
FP8 GEMM · FP8 matrix multiplication (GEMM) evaluated · 1 source · 1 quote
Avoiding split-K · we abandon split-k in most scenarios not used · 1 source · 1 quote

context management 18

Preserved thinking history mode · preserved thinking, Thinking Preservation core · 7 sources · 7 quotes
Excluding prior thinking from conversation history · No Thinking Content in History default · 6 sources · 6 quotes
Discard-all context management · Discard-all, discard-all context management strategy used · 6 sources · 7 quotes
Context compaction · context compaction at 300K tokens, Context-compaction strategy used · 3 sources · 3 quotes
Context folding · context-folding strategy, simple context-folding strategy(256k) used · 2 sources · 2 quotes
Thinking context management for tool use · Thinking Context Management used · 2 sources · 2 quotes
Context management method used · 1 source · 1 quote
Context management strategy used · 1 source · 1 quote
Discarding tool-call history used · 1 source · 1 quote
Hierarchical context management · Hierarchical Context Management strategy used · 1 source · 1 quote
Keep-recent-k · keep-recent-k context management, Keep-recent-k strategy used · 1 source · 1 quote
Memory compression · memory compression (context consolidation) used · 1 source · 1 quote
Preserve thinking · Preserve thinking blocks optional · 2 sources · 2 quotes
Modality-specific deployment · deploying only the modalities you need optional · 1 source · 1 quote
Discard-75% evaluated · 2 sources · 2 quotes
Trajectory summarization and rollout re-initiation · Summary evaluated · 2 sources · 2 quotes
Summary-based context compression · summary-based compression evaluated · 1 source · 1 quote

agentic scaffolding 93

Function calling · Function calling (structured tool use), native function calling core · 3 sources · 3 quotes
Harmony format · harmony response format core · 3 sources · 3 quotes
Role-based instruction hierarchy · role hierarchy core · 2 sources · 2 quotes
Assistant output channels · channels core · 1 source · 1 quote
Chain-of-thought in the analysis channel core · 1 source · 1 quote
Developer message format core · 1 source · 1 quote
Function-calling format core · 1 source · 1 quote
Harmony channel annotations core · 1 source · 1 quote
Harmony tool-call message format core · 1 source · 1 quote
Sandbox snapshots · Snapshot core · 1 source · 1 quote
Sandbox state forking · Fork core · 1 source · 1 quote
Automatic tool choice · auto-tool choice, enable-auto-tool-choice used · 4 sources · 4 quotes
Python tool used · 4 sources · 4 quotes
Browser tool · web browsing tool training used · 2 sources · 2 quotes
Generic tool-call parser · Tool call parser used · 2 sources · 2 quotes
Multi-agent collaboration used · 2 sources · 2 quotes
Visual Search Tool used · 2 sources · 2 quotes
XML-based tool-call schema with DSML token · XML-based tool-call schema used · 2 sources · 2 quotes
Agent Team mode used · 1 source · 1 quote
AgentENV used · 1 source · 1 quote
Agentic search used · 1 source · 1 quote
App-server mode with adapted tool schemas used · 1 source · 1 quote
AppArmor and eBPF sandbox policies used · 1 source · 1 quote
Asynchronous teammate spawning · spawn_teammate used · 1 source · 1 quote
Bash commands for context retrieval · Bash commands used · 1 source · 1 quote
Bash computer-use agent used · 1 source · 1 quote
Browsing tool with domain filtering · browsing tool with a domain block used · 1 source · 1 quote
Commentary-channel preambles · preambles on the commentary channel used · 1 source · 1 quote
Concurrent subagent orchestration · using more than 20 concurrent subagents used · 1 source · 1 quote
Current-turn-only reasoning-mode detection · scans only the current turn’s content used · 1 source · 1 quote
Developer-defined function schemas used · 1 source · 1 quote
Dynamic tool loading · dynamically loaded tools used · 1 source · 1 quote
Fresh and fork teammate initialization modes · fresh mode or fork mode used · 1 source · 1 quote
Harmony history stop-token normalization used · 1 source · 1 quote
Hy v4 tool-call parser · Tool-call parser hy_v4 used · 1 source · 1 quote
Indexed parallel tool calls · tool and index attributes used · 1 source · 1 quote
Interleaving tool calls with chain-of-thought · interleaving tool calls within the CoT used · 1 source · 1 quote
JSON Schema response formats · structured output response formats used · 1 source · 1 quote
Jupyter Notebook code interpreter · Jupyter Notebook as a code interpreter used · 1 source · 1 quote
Lead-agent interruption of teammates · Lead-agent interruption of teammate turns, interrupt_agent used · 1 source · 1 quote
Mini-harnesses used · 1 source · 1 quote
Native task delegation used · 1 source · 1 quote
Programmatic TypeScript tool calling used · 1 source · 1 quote
Prompt-based cold start for reasoning in tool use · Cold-Start used · 1 source · 1 quote
Prompt-enforced tool-call format · our designed toolcall format used · 1 source · 1 quote
Qwen3 XML tool-call parser · tool-call parser qwen3_xml used · 1 source · 1 quote
ReAct Toolbelt · ReAct Toolbelt framework used · 1 source · 1 quote
Request timeouts and retries for agent loops · set a request timeout and a retry used · 1 source · 1 quote
Retrieval-Augmented Search used · 1 source · 1 quote
Runtime command filter used · 1 source · 1 quote
Sandbox infrastructure (DSec) used · 1 source · 1 quote
Sandboxed Python execution loop · secure sandboxed Python execution loop used · 1 source · 1 quote
Scrollable browser text window · scrollable window of text used · 1 source · 1 quote
Shared task board with revision checks used · 1 source · 1 quote
Single-Bash-tool scaffold interface · singlebash tool used · 1 source · 1 quote
Stateful Python tool · stateful tool used · 1 source · 1 quote
Stateless Python tool reference implementation · stateless mode used · 1 source · 1 quote
Step3p5 tool-call parser · tool-call parser step3p5 used · 1 source · 1 quote
Streaming-delta handling across block boundaries · Deltas spanning block boundaries used · 1 source · 1 quote
System message format used · 1 source · 1 quote
Thinking with tools · Thinking with tools capability used · 1 source · 2 quotes
Tool calls within the thinking process used · 1 source · 1 quote
Tool output message format · tool message format used · 1 source · 1 quote
Tool-Integrated Reasoning · TIR used · 1 source · 1 quote
Tools section in the system message · # Tools section used · 1 source · 1 quote
Typed tool arguments · Arguments are typed used · 1 source · 1 quote
TypeScript-like function schema syntax · TypeScript-like type syntax for functions used · 1 source · 1 quote
Vision in the loop used · 1 source · 1 quote
Web search extension used · 1 source · 1 quote
XML-tagged tool-call format used · 1 source · 1 quote
XTML chat template · XTML-based chat template used · 1 source · 1 quote
Advisor strategy · Advisor Mode optional · 2 sources · 4 quotes
MCP tool configuration · MCP configuration file for available tools, Tool use via MCP configuration optional · 2 sources · 2 quotes
Checkpointing · Checkpointing for rollback of agent edits, checkpointing feature optional · 1 source · 1 quote
Claude Code skills for Tinker · Claude Code skills optional · 1 source · 1 quote
Compact TXT image-path notation · compact TXT image-path prompt encoding, compact <image>path</image> TXT notation optional · 1 source · 1 quote
JSON mode optional · 1 source · 1 quote
OpenAI-style JSON content blocks optional · 1 source · 1 quote
Qwen3 Coder tool-call parser · --tool-call-parser qwen3_coder optional · 1 source · 1 quote
StreamableParser optional · 1 source · 1 quote
Tool calling optional · 1 source · 1 quote
Video understanding and editing optional · 1 source · 1 quote
Vision-driven UI coding optional · 1 source · 1 quote
Agentic tool calling mentioned · 1 source · 1 quote
Task-specification prompting mentioned · 1 source · 1 quote

software implementation 79

filed at the root 7

Separate encoding and inference modules · separate encoding/ and inference/ default · 1 source · 1 quote
Data Designer used · 2 sources · 2 quotes
CUDA used · 1 source · 1 quote
fastsafetensors · fastsafetensors checkpoint loading, fastsafetensors load format used · 1 source · 1 quote
Nemotron Parse used · 1 source · 1 quote
Model routing and orchestration optional · 1 source · 1 quote

inference engine 15

vLLM · Serving MiMo-V2.6-Pro-RL with vLLM, vLLM model serving used · 7 sources · 7 quotes
TensorRT-LLM · TRTLLM, TensorRT-LLM inference optimization used · 2 sources · 3 quotes
Atlas inference library · Atlas inference library on vLLM used · 1 source · 1 quote
Environment variable propagation to subprocesses · Environment propagation used · 1 source · 1 quote
litert-lm serve used · 1 source · 1 quote
llama.cpp · llama.cpp inference library used · 1 source · 1 quote
vLLM or SGLang serving · vLLM/SGLang serving, vLLM or SGLang used · 1 source · 1 quote
vLLM tensor parallelism · vLLM at --tensor-parallel-size 8 used · 1 source · 1 quote
SGLang · Serving MiMo-V2.6-Pro-RL with SGLang optional · 4 sources · 4 quotes
KTransformers optional · 2 sources · 2 quotes
Unsloth optional · 2 sources · 2 quotes
Dedicated serving engines optional · 1 source · 1 quote
vLLM language-model-only mode · --language-model-only optional · 1 source · 1 quote
vLLM-Ascend optional · 1 source · 1 quote
xLLM optional · 1 source · 1 quote

training framework 6

DAG-based asset lineage tracking · directed acyclic graph (DAG) of assets core · 1 source · 1 quote
Experiments as Code core · 1 source · 1 quote
Hugging Face Transformers · Transformers, HuggingFace Transformers used · 4 sources · 4 quotes
Megatron-LM used · 3 sources · 3 quotes
NeMo Gym used · 2 sources · 2 quotes
NeMo · NeMo 26.04.01 used · 1 source · 1 quote

kernel & quantization library 18

FlashInfer · FlashInfer backend used · 3 sources · 3 quotes
AngelSlim · AngelSlim compression toolkit, AngelSlim toolkit used · 2 sources · 2 quotes
FlashInfer NVLinkOneSided · FlashInfer NVLinkOneSided All-to-All, FlashInfer's NVLinkOneSided used · 2 sources · 2 quotes
FlashInfer TensorRT-LLM backend · flashinfer_trtllm used · 2 sources · 2 quotes
MiniTriton used · 2 sources · 2 quotes
NVIDIA Model-Optimizer · Model-Optimizer used · 2 sources · 2 quotes
AITER · AITER linear and MoE backends, AITER backend used · 1 source · 1 quote
Container-baked precompiled artifacts · Container-baked artifacts used · 1 source · 1 quote
FLASHMLA_SPARSE · FLASHMLA_SPARSE attention backend used · 1 source · 1 quote
FlashQLA used · 1 source · 1 quote
Marlin MoE backend · moe-backend marlin used · 1 source · 1 quote
MiniTriton dual-mode tensor library used · 1 source · 1 quote
CUTEDSL optional · 1 source · 1 quote
CUTLASS optional · 1 source · 1 quote
deepseek-recipe optional · 1 source · 1 quote
flash-moba evaluated · 1 source · 1 quote
Flash-Sparse-Attention evaluated · 1 source · 1 quote
Triton mentioned · 1 source · 1 quote

agent product 11

OpenCode used · 4 sources · 4 quotes
Claude Code used · 3 sources · 3 quotes
OpenHands used · 2 sources · 2 quotes
Qwen-Agent used · 2 sources · 2 quotes
Droid used · 1 source · 1 quote
Mini-SWE-Agent used · 1 source · 1 quote
OpenClaw · OpenClaw technology used · 1 source · 1 quote
Pi used · 1 source · 1 quote
Stirrup used · 1 source · 1 quote
Kilo optional · 1 source · 1 quote

infrastructure service 22

Automatic provider failover · automatic retry on next-best provider, provider failover routing core · 6 sources · 6 quotes
DeepSeek Elastic Compute (DSec) core · 1 source · 1 quote
End-to-end lineage core · 1 source · 1 quote
Sandbox pause and resume · pause-and-resume sandbox lifecycle control, Pause and Resume core · 1 source · 1 quote
Custom placement engine used · 1 source · 1 quote
Dependency unification used · 1 source · 1 quote
Durable peer mailbox · Durable peer mailbox communication used · 1 source · 1 quote
Fresh-clone repository sanitization used · 1 source · 1 quote
HTTP/HTTPS sidecar proxy used · 1 source · 1 quote
In-house batch scheduler · custom cluster scheduler used · 1 source · 1 quote
Local container-image caching · Container caching used · 1 source · 1 quote
Model Factory used · 1 source · 1 quote
Mooncake used · 1 source · 1 quote
NeMo Guardrails · NeMo Guardrails policy enforcement used · 1 source · 1 quote
NeMo Switchyard · NeMo Switchyard model routing used · 1 source · 1 quote
Per-node pooling of initialization actors · pooling initialization actors per node used · 1 source · 1 quote
Replacing short-lived Ray actors with tasks · converting short-lived actors to tasks used · 1 source · 1 quote
Vendoring external dependencies used · 1 source · 1 quote
Write-ahead log used · 1 source · 1 quote
Z3 SMT solver integration used · 1 source · 1 quote
Ray optional · 1 source · 1 quote

evaluation 109

filed at the root 35

Prompt-based output standardization · Standardize Output Format, standardize model outputs used · 5 sources · 5 quotes
Evaluation without safety filters · safety evaluation without safety filters, safety evaluations without safety filters used · 3 sources · 3 quotes
Refusal-suppressed variants for capability estimation · refusal-suppressed variants used · 3 sources · 3 quotes
avg@k · average over 3 rollouts (avg@3), avg@3 used · 2 sources · 2 quotes
Almost@1 · Almost@1 rollout success metric used · 1 source · 1 quote
Behavior testing with stubs and edge cases · behaviour test for every bash script used · 1 source · 1 quote
Deadline-bounded rollout evaluation · explicit per-rollout wall-clock deadlines used · 1 source · 1 quote
Documenting evaluation configuration · document our exact settings used · 1 source · 1 quote
End-to-end exploit development evaluation · Exploit development (Tier 2) used · 1 source · 1 quote
Evaluation equivalence threshold · Evaluation equivalence threshold of 0.3 used · 1 source · 1 quote
Fixed MTP draft-token acceptance length · fixed acceptance length used · 1 source · 1 quote
Four-run mean pass@1 · Mean pass@1 across four runs used · 1 source · 1 quote
Full trajectory release used · 1 source · 1 quote
Joint validation protocol · joint protocol used · 1 source · 1 quote
Mean@5 · Mean@5 evaluation metric used · 1 source · 1 quote
Prompt engineering to elicit answers · enforce an answer via prompt engineering used · 1 source · 1 quote
ProtocolQA robustness validation · ProtocolQA robustness checks used · 1 source · 1 quote
Refusal behavior quantification · Quantified refusal behavior used · 1 source · 1 quote
Rollout-based auditing used · 1 source · 1 quote
Standard function-calling format · standard function call format used · 1 source · 1 quote
Standardized agent evaluation protocol used · 1 source · 1 quote
Trajectory analysis used · 1 source · 1 quote
Two-dimensional coding-agent evaluation used · 1 source · 1 quote
Union-of-capabilities coverage scoring · union of capabilities across revisions used · 1 source · 1 quote
Visual checking of generated renders · visual checks used · 1 source · 1 quote
Vulnerability discovery with proof-of-concept evaluation · Vulnerability discovery (Tier 1) used · 1 source · 1 quote
Zero-shot base-model baseline fallback used · 1 source · 1 quote
Model–harness co-design optional · 1 source · 1 quote
V-model iterative testing and validation · V-model methodology mentioned · 1 source · 1 quote

benchmark 25

Living in-house benchmark suite core · 1 source · 1 quote
BrowseComp with Search used · 1 source · 1 quote
Capture-the-Flag (CTF) evaluation · CTF challenges used · 1 source · 1 quote
Disabling tool search · disabling tool search in evaluation, Tool Search disabled used · 1 source · 1 quote
Domain whitelist · domain whitelist for agent environments used · 1 source · 1 quote
Execution-grounded evaluation used · 1 source · 1 quote
F2P/P2F acceptance criterion used · 1 source · 1 quote
Fail-to-pass and pass-to-pass evaluation points · fail-to-pass and pass-to-pass points used · 1 source · 1 quote
GPU kernel optimization task suite · kernel optimization tasks used · 1 source · 1 quote
High-confidence ProgramBench subset · High-confidence benchmark subset filtering, high-confidence subset of ProgramBench used · 1 source · 2 quotes
KernelBench used · 1 source · 1 quote
LongBench v2 used · 1 source · 1 quote
Multiple-rollout evaluation · Multiple-rollout evaluation per task, up to three rollouts per task used · 1 source · 1 quote
PaperBench used · 1 source · 1 quote
PinchBench used · 1 source · 1 quote
ProfBench with Search used · 1 source · 1 quote
SQuAD used · 1 source · 1 quote
SWE-bench Verified used · 1 source · 1 quote
Three-configuration cyber range testing used · 1 source · 1 quote
Vals.ai used · 1 source · 1 quote
Verified CUDA kernels · Verified CUDA kernels for synthetic data used · 1 source · 1 quote
CAD visual reproduction optional · 1 source · 1 quote
Computer use closed-loop evaluation · Computer use closed loop optional · 1 source · 1 quote
Artificial Analysis Intelligence Index evaluated · 1 source · 1 quote

evaluation harness 14

DeepSeek Harness Minimal mode · DeepSeek Harness Minimal evaluation mode, Minimal mode of DeepSeek Harness used · 2 sources · 2 quotes
Harbor used · 2 sources · 2 quotes
NeMo Evaluator SDK used · 2 sources · 2 quotes
Atomic binary rubric evaluation used · 1 source · 1 quote
Benchmark evaluation framework · benchmark framework used · 1 source · 1 quote
Cross-scaffold evaluation · Cross-scaffold robustness evaluation used · 1 source · 1 quote
Isolated container per rollout · isolated container per agent rollout used · 1 source · 1 quote
JSON-structured answer-format prompt used · 1 source · 1 quote
mini-swe-agent harness used · 1 source · 1 quote
NeMo Gym and NeMo Evaluator-based harness used · 1 source · 1 quote
NeMo Skills used · 1 source · 1 quote
Pool · agentic coding evaluations run using pool used · 1 source · 1 quote
Inline training evaluators optional · 1 source · 1 quote

judge 16

Agent-as-a-Judge · Agent-as-a-Judge evaluation method used · 1 source · 2 quotes
Agent-as-a-Verifier used · 1 source · 1 quote
Decompositional instruction-following judge · dedicated IF judge used · 1 source · 1 quote
Deterministic tool-call verifier · deterministic verifier used · 1 source · 1 quote
GPT-5.5 (medium) judge model · GPT-5.5 (medium) as the judge model used · 1 source · 1 quote
Human-calibrated LLM-as-a-Judge used · 1 source · 1 quote
Independent quality-inspection agent used · 1 source · 2 quotes
LLM autograding with expert validation used · 1 source · 1 quote
LLM-as-a-Judge used · 1 source · 1 quote
LLM-based anti-cheating judgment · LLM-based judgement for anti-cheating, LLM-based judgement used · 1 source · 1 quote
LLM-based correctness judging · LLM-based judging pipeline used · 1 source · 1 quote
LLM-based external API usage inspection used · 1 source · 1 quote
Mandatory agentic judge protocol · mandatory protocol for agentic judge used · 1 source · 1 quote
Official task verifier scoring · official separate verifier used · 1 source · 1 quote
Reward-hack detection judge · Reward hack judge used · 1 source · 1 quote
Rule-based anti-cheat checks · rule-based judgement used · 1 source · 1 quote

human & real-world evaluation 19

Held-out evaluation · held-out generalization gates default · 1 source · 1 quote
Blind side-by-side expert evaluation · blind expert judging of model outputs, blind expert judging used · 4 sources · 4 quotes
Multi-turn open-ended external red-teaming · Multi-turn open-ended red-teaming used · 3 sources · 3 quotes
Automated red-teaming loop used · 1 source · 1 quote
Cyber range exercises used · 1 source · 1 quote
Dangerous-capability uplift assessment used · 1 source · 1 quote
Human evaluation · human evaluations used · 1 source · 1 quote
Internal engineering task evaluation · Internal engineering tasks evaluation used · 1 source · 2 quotes
Pairwise comparison for Chinese writing · Pairwise comparisons for Chinese writing used · 1 source · 1 quote
Pre-release safety evaluation used · 1 source · 1 quote
Real-world codebase testing used · 1 source · 1 quote
Real-world feedback refinement used · 1 source · 1 quote
Reward hacking mitigation in evaluations · reward hacking mitigation used · 1 source · 1 quote
Same-prompt comparative DevOps testing · same-prompt DevOps test used · 1 source · 1 quote
White-collar enterprise task evaluation · White-collar enterprise tasks evaluation used · 1 source · 1 quote
A/B testing on real tasks · A/B testing on three real tasks, A/B on three real tasks mentioned · 1 source · 1 quote
Continuous monitoring mentioned · 1 source · 1 quote

other 16

filed at the root 16

Automating Repetitive Infrastructure Work core · 1 source · 1 quote
Agentic Workflows used · 1 source · 1 quote
Code-Driven Animation used · 1 source · 1 quote
Early Release with Feedback-Driven Iteration · ship early and hear what breaks used · 1 source · 1 quote
Prompt Modality Ordering · Modality order used · 1 source · 1 quote
Removing Git-History Leaks from Benchmark Images · removed to prevent possible reward hacking used · 1 source · 2 quotes
Request Caching in the Browser Tool · the tool caches requests used · 1 source · 1 quote
Reward-Hacking Detection System · hacking-detection system used · 1 source · 1 quote
Score-versus-Cost Efficiency Comparison · comparing score against per-task cost used · 1 source · 1 quote
Training–Inference Consistency · train–inference consistency used · 1 source · 1 quote
Unique Run Identifiers · unique ID used · 1 source · 1 quote
Closed-Model Performance with Refusal Substitution · combined performance using closed models not used · 1 source · 1 quote
Removing Pattern-Matching-Based Anti-Cheat Checks · removing pattern-matching-based checks not used · 1 source · 1 quote

unfiled 231

Gated Attention core · 3 sources · 3 quotes
Gated DeltaNet and Qwen Sparse Attention hybrid architecture · GDN + QSA hybrid architecture core · 2 sources · 2 quotes
Adaptive Rate Interleave Alignment (ARIA) · ARIA (Adaptive Rate Interleave Alignment) core · 1 source · 1 quote
Agentic capabilities core · 1 source · 1 quote
agentic model core · 1 source · 2 quotes
Blockwise candidate-pool selection for sparse indexing · blockwise candidate selection core · 1 source · 1 quote
disk-based context caching · hard disk caching core · 1 source · 1 quote
DSpark forward path core · 1 source · 1 quote
Gated DeltaNet and global-attention hybrid core · 1 source · 1 quote
Gated DeltaNet and Qwen Sparse Attention hybrid · GDN + QSA core · 1 source · 1 quote
Global attention with a dense feed-forward network in the first Transformer block · global attention with a dense FFN core · 1 source · 1 quote
harmony chat format core · 1 source · 1 quote
harness awareness core · 1 source · 1 quote
Hybrid fast-and-slow-thinking model · hybrid fast-and-slow-thinking core · 1 source · 1 quote
hybrid reasoning architecture core · 1 source · 1 quote
IndexShare for multi-token prediction speculative decoding · IndexShare MTP core · 1 source · 1 quote
Interleaved multimodal input core · 1 source · 1 quote
Micro-block compressed lightweight indexer core · 1 source · 1 quote
Model specialization for agent execution core · 1 source · 1 quote
multimodal training from step zero core · 1 source · 1 quote
MXFP4 weight storage with block size 32 · block size of 32 core · 1 source · 1 quote
native multimodal training from the start of training · native multimodal training core · 1 source · 1 quote
Native system prompt support core · 1 source · 1 quote
OpenAI Harmony response format core · 1 source · 1 quote
post-training core · 1 source · 1 quote
post-training scaling · scaled post-training core · 1 source · 1 quote
pre-training scaling core · 1 source · 1 quote
scaling post-training core · 1 source · 1 quote
Short convolution (sconv) modules · Short convolution (sconv) core · 1 source · 1 quote
Transformer core · 1 source · 1 quote
OpenAI-compatible API · OpenAI-compatible default · 3 sources · 3 quotes
configurable video frame sampling with fps and do_sample_frames · fps=2 and do_sample_frames=True default · 1 source · 1 quote
Omit residual branch-mixing operator · removing the mixing operator default · 1 source · 1 quote
PLE CPU offload for N-gram embeddings · PLE CPU offload default · 1 source · 1 quote
Terminus 2 · Terminus-2 framework used · 2 sources · 2 quotes
30T-token multimodal pre-training corpus used · 1 source · 1 quote
adequate output length (32,768 tokens default, 81,920 for hard benchmarks) · Adequate Output Length used · 1 source · 1 quote
agentic repository installation task used · 1 source · 1 quote
Amazon Elastic Container Service · AWS ECS used · 1 source · 1 quote
Apache 2.0 License · Apache 2.0 used · 1 source · 1 quote
archipelago · archipelago codebase used · 1 source · 1 quote
architectural innovation used · 1 source · 1 quote
autonomous inference system optimization used · 1 source · 1 quote
benchmarking against public frontier models · benchmarked against public frontier models used · 1 source · 1 quote
benchmaxxing used · 1 source · 1 quote
chat templates for model prompting · model chat templates used · 1 source · 1 quote
Claude Code harness used · 1 source · 1 quote
Configurable routing profiles for model selection · configurable routing profiles used · 1 source · 1 quote
contrastive pre-training of vision encoder · contrastive pre-training used · 1 source · 1 quote
cross-lingual voice cloning used · 1 source · 1 quote
cross-provider failover load balancing · routing to another healthy provider used · 1 source · 1 quote
Cross-provider failover routing for inference requests · routing to another healthy provider used · 1 source · 1 quote
custom-voice speech generation used · 1 source · 1 quote
Data quality and diversity enhancement · enhanced data quality and diversity used · 1 source · 1 quote
efficient clustering in the embedder used · 1 source · 1 quote
Failover routing across providers · routing to another healthy provider used · 1 source · 1 quote
Failover routing to healthy upstream providers · routing to another healthy provider used · 1 source · 1 quote
fine-grained EP scheme used · 1 source · 1 quote
Fine-tuning hyperparameter verification · Fine-tuning performance verification used · 1 source · 1 quote
Fine-tuning via Tinker platform · fine-tune themselves through Tinker used · 1 source · 1 quote
Fixing dependency drift in benchmark tasks · Fixed dependency drift on tasks used · 1 source · 1 quote
Fixing verifier selections in benchmark tasks · fixing verifier selections used · 1 source · 1 quote
FlashComm used · 1 source · 1 quote
four-stage Talker training pipeline · four-stage training pipeline for Talker used · 1 source · 1 quote
frontier/execution model division of labor used · 1 source · 1 quote
Fuse gated residual reads and writes into single kernels · fused into a single kernel used · 1 source · 1 quote
GPT-OSS-120B used · 1 source · 1 quote
GUI operation used · 1 source · 1 quote
HelpSteer3 used · 1 source · 1 quote
HotpotQA used · 1 source · 1 quote
Ignoring known teardown flakes in evaluation · we ignore it used · 1 source · 1 quote
In-task retries for flaky external services · in-task retries used · 1 source · 1 quote
Inference-time scaling plots for evaluation reporting · Inference-time scaling plots used · 1 source · 1 quote
JIT cache management used · 1 source · 1 quote
Large collection of reinforcement-learning task environments · more than 7,000 RL task environments used · 1 source · 1 quote
likelihood-based acceptance-rate loss (LK loss) · LK loss for draft model fine-tuning used · 1 source · 1 quote
Limit maximum concurrent sequences to 256 · keep --max-num-seqs 256 used · 1 source · 1 quote
Live reinforcement-learning training used · 1 source · 1 quote
LSE fusion used · 1 source · 1 quote
masking-based refinement · mask fine-tuning used · 1 source · 1 quote
Max-pool teacher attention distributions to block level · max pooling used · 1 source · 1 quote
Measure per-block maximum activation · per-block maximum activation used · 1 source · 1 quote
MLAPO used · 1 source · 1 quote
multilingual speech generation used · 1 source · 1 quote
multimodal pre-training · multimodal pre-training at scale used · 1 source · 1 quote
multimodal prompt content ordering used · 1 source · 1 quote
ngspice simulation · ngspice simulation loop used · 1 source · 1 quote
NVIDIA API Trial Terms of Service used · 1 source · 1 quote
Open Model Data Warehouse License Agreement 1.1 · OpenMDW License Agreement, version 1.1 used · 1 source · 1 quote
open-weight model release · open-weight systems used · 1 source · 1 quote
Open-weights release under Apache-2.0 license · open-weights model used · 1 source · 1 quote
OpenAI-compatible API endpoint · OpenAI-compatible endpoint used · 1 source · 1 quote
OpenAI-compatible message encoding used · 1 source · 1 quote
OpenAI-compatible message encoding without Jinja template · encoding folder used · 1 source · 1 quote
Optimization on real agent-harness environments · Optimized on real harness environments used · 1 source · 1 quote
pass@1 metric · pass@1 used · 1 source · 1 quote
Peak/off-peak pricing used · 1 source · 1 quote
Per-cell power-law fits for loss vs tokens used · 1 source · 1 quote
per-model token-per-second rescaling of timeouts · per-model TPS rescaling used · 1 source · 1 quote
Pinning dependency versions to avoid drift · pinned to setuptools^=58.0.0 used · 1 source · 1 quote
Post-training pipeline used · 1 source · 1 quote
Pre-commit hooks with ruff · pre-commit hooks used · 1 source · 1 quote
pre-training and post-training used · 1 source · 1 quote
Pre-training from scratch · pre-trained Inkling from scratch used · 1 source · 1 quote
preview-first release approach · preview-first approach used · 1 source · 1 quote
private benchmark for evaluation · private benchmark used · 1 source · 1 quote
provider routing modes used · 1 source · 1 quote
Publication of unedited evaluation trajectories · three unedited trajectories used · 1 source · 1 quote
Python (stateful Jupyter) tool training used · 1 source · 1 quote
query concatenation used · 1 source · 1 quote
Qwen3 reasoning parser · reasoning parser qwen3 used · 1 source · 1 quote
Ralph-Loop mechanism used · 1 source · 1 quote
Reasoning parser hy_v4 used · 1 source · 1 quote
recommended sampling parameters · sampling parameters used · 1 source · 1 quote
recommended sampling parameters for deployment · sampling parameters used · 1 source · 1 quote
recommended sampling parameters for thinking and non-thinking modes · Sampling Parameters used · 1 source · 1 quote
recursive self-improvement loop used · 1 source · 1 quote
Reproducible evaluation recipes in NeMo Gym · evaluation recipes used · 1 source · 1 quote
Request-filter-constrained provider routing · request filters used · 1 source · 1 quote
Rolling-median loss-spike measurement used · 1 source · 1 quote
Run ShellCheck at info severity to catch possible misspellings · Run ShellCheck at info level used · 1 source · 1 quote
shadow replicas for shared indexers in pipeline-parallel training · Shadow indexers used · 1 source · 1 quote
shared-memory caching of preprocessed multimodal inputs · --mm-processor-cache-type shm used · 1 source · 1 quote
Single multi-node Slurm launch · single multi-node srun used · 1 source · 1 quote
staged release · staged approach to release used · 1 source · 1 quote
Tau Bench 3 used · 1 source · 1 quote
Terminus used · 1 source · 1 quote
Text-only multimodal benchmark alignment for comparability · Multimodal benchmark alignment used · 1 source · 1 quote
three-stage SWE teacher training pipeline · three-stage pipeline used · 1 source · 1 quote
TileLang used · 1 source · 1 quote
tool augmentation during evaluation · tool augmentation used · 1 source · 1 quote
tool-augmented benchmark evaluation (with and without tools) · tool augmentation used · 1 source · 1 quote
Training to resist censorship used · 1 source · 1 quote
Two-stage contextual-parallel communication for compressed attention · two-stage communication approach used · 1 source · 1 quote
WSD optimal learning rate scaling law · WSD optimal-LR scaling law used · 1 source · 1 quote
cross-provider failover routing · routing to another healthy provider optional · 2 sources · 2 quotes
3D scene creation with Blender optional · 1 source · 1 quote
Adequate output length recommendation · output length of 32,768 tokens optional · 1 source · 1 quote
AITER Linear backend · VLLM_ROCM_USE_AITER_LINEAR=1 optional · 1 source · 1 quote
AITER MHA backend · VLLM_ROCM_USE_AITER_MHA=1 optional · 1 source · 1 quote
AITER RMSNorm backend · VLLM_ROCM_USE_AITER_RMSNORM=1 optional · 1 source · 1 quote
Balanced routing: price + speed per request · Balanced (price + speed) optional · 1 source · 1 quote
Configurable video frame sampling (fps/do_sample_frames) during inference · video frame sampling optional · 1 source · 1 quote
Financial analysis workflow optional · 1 source · 1 quote
Fine-tuning on Tinker platform · availability on Tinker for fine-tuning optional · 1 source · 1 quote
Game development workflow optional · 1 source · 1 quote
Multi-precision weight formats (BF16/FP8/INT4/NVFP4) · weights in BF16, FP8, INT4, and NVFP4 optional · 1 source · 1 quote
Nitro routing: fastest provider per request · Nitro (fastest) optional · 1 source · 1 quote
Office deliverables creation optional · 1 source · 1 quote
OpenAI Harmony renderer (openai_harmony) that renders messages into tokens · harmony renderer library optional · 1 source · 1 quote
Output length recommendation · output length of 32,768 tokens optional · 1 source · 1 quote
Recommended API parameter settings · Recommended Settings optional · 1 source · 1 quote
reducing CUDA graph capture size to fit Mamba cache · Reduce --max-cudagraph-capture-size optional · 1 source · 1 quote
request routing modes · routing mode optional · 1 source · 1 quote
Serving Qwen3.5 with SGLang · SGLang serving framework optional · 1 source · 1 quote
Serving Qwen3.5 with vLLM · vLLM serving engine optional · 1 source · 1 quote
task-to-best-model routing with NeMo Switchyard optional · 1 source · 1 quote
Interaction models (AI that listens, speaks, and interrupts) · interaction models evaluated · 1 source · 1 quote
self-verification evaluated · 1 source · 1 quote
Tiered CPU and filesystem offloading · Validation used TieringOffloadingSpec evaluated · 1 source · 1 quote
Bias audits mentioned · 1 source · 1 quote
buying benchmark-specific training data mentioned · 1 source · 1 quote
cache-hit cost modeling for agent workloads · model the cache-hit price first mentioned · 1 source · 1 quote
Defense-in-depth for safety · defense-in-depth mentioned · 1 source · 1 quote
Dynamic Workflows mentioned · 1 source · 1 quote
Emergent Compositional Tool Use mentioned · 1 source · 1 quote
Incremental task decomposition for agent prompts mentioned · 1 source · 1 quote
Input/output classification with moderation tools · input/output classification mentioned · 1 source · 1 quote
Planning with the model for self-contained tasks · planning with the model mentioned · 1 source · 1 quote
Pre-training methods · New pre-training methods mentioned · 1 source · 1 quote
Sampling with temperature=1.0 and top_p=1.0 · temperature=1.0 and top_p=1.0 mentioned · 1 source · 1 quote
Serving Qwen3.5 with KTransformers mentioned · 1 source · 1 quote
AITER fused-MoE backend · VLLM_ROCM_USE_AITER_MOE=1 not used · 1 source · 1 quote
Read only the two highest-gated residual branches · sparse writes not used · 1 source · 1 quote