1 733 methods · 1 502 filed · 2 955 quotes
Every method, filed where it belongs.
1 733 methods extracted from 164 documents, nested the way the curated taxonomy files them. Within each group, the strongest treatment any source gives a method comes first.
model architecture 312
filed at the root 7
Causal Encoder-Decoder (CED) architecture · Causal Encoder-Decoder architecture, Causal Encoder-Decoder (CED)
Per-Layer Embeddings (PLE) · per-layer embeddings
Encoder-free decoder-only Transformer architecture · encoder-free architecture
Dense model architecture · 31B Dense
token mixer 123
Cross-layer attention sharing · attention sharing method
Micro-batch-level shared-state lifetime management · Micro-batch-level shared-state management
Dense gated attention with full RoPE and full gating · Dense GA, full RoPE, full gating
softmax attention 32
global attention 4
Global attention · Global Attention (GA)
Unified Keys and Values · unified KV for global attention layers
Multi-Head Attention (MHA) · Multi-Head Attention (MHA) for reliability
sliding window attention 12
Sliding Window Attention · SWA, Sliding Window Attention (SWA)
Hybrid Sliding Window Attention · Hybrid window attention
Sink-Augmented Sliding Window Attention · sink-augmented SWA
512-token Sliding Window Attention · + SWA-512
Decoder SWA Bounded Replay · decoder bounded replay during prefill
SWA-128 · Sliding Window Attention with size 128
Fixed Sliding Window Attention · sliding-window baseline
Per-Head Gating · Per-head gating with θswa=10,000, + per-head gating, θswa = 10,000
SWA-1024 · + SWA-1024 (interleaved 3:1)
grouped-query attention 4
Grouped-query attention · GQA, GQA, 8 KV heads
Block-sparse grouped-query attention · block-sparse GQA
Multi-query attention · Multi-Query Attention (MQA), Multi-Query Attention (MQA) mode of MLA
Softplus-based per-head gating · Softplus-based per-head attention gating
multi-head latent attention 4
Gated Multi-Head Latent Attention · Gated MLA
MQA Mode of Multi-Head Latent Attention · MQA mode of MLA
Multi-Head Latent Attention · MLA, Multi-latent Attention
Multi-Head Latent Attention MHA/MQA Modes · MHA and MQA modes of MLA
attention sink 3
Learnable attention sink bias · Attention sink bias, learnable attention sink
RoPE with attention sink · RoPE + attention sink
sparse attention 57
DeepSeek Sparse Attention · DSA, DSA (DeepSeek Sparse Attention)
Compressed Sparse Attention · CSA, Compressed Sparse Attention (CSA)
Qwen Sparse Attention · Qwen Sparse Attention (QSA)
Sparse attention · sparse attention layers
Heavily Compressed Attention · Heavily Compressed Attention (HCA)
Fixed-budget sparse-attention selection · fixed budget of 512 blocks or 2048 tokens
NoPE sparse multi-head latent attention · NoPE sparse MLA layers
KV-outer sparse attention · KV outer gather Q
Sparse-attention continued pre-training with joint model and indexer optimization · Sparse Training Stage
CSA2 Full Mode · Full Mode
Joint backbone and indexer training under sparse attention · sparse training
Sparse softmax attention · sparse softmax attention paradigm
Two-stage introduction of sparse attention · two-stage training method
Two-stage sparse attention · two-stage sparse-attention formulation
Native Sparse Attention · NSA
block selection 9
MiniMax Sparse Attention · MSA (MiniMax Sparse Attention), MiniMax Sparse Attention (MSA)
Top-512 block selection · QSA keeps the best 512 blocks
Top-k block selection · Top- block selection, top-k selection
Block Max Pooling · Block Max Pool
Mixture of Block Attention · MoBA
sparse attention indexer 22
Compressed Sparse Attention 2 · Compressed Sparse Attention 2 (CSA2), cross-layer KV-cache and index reuse
Fine-grained token selection · fine-grained token selection mechanism
Hierarchical Sparse Indexer · Hierarchical sparse indexing
Reindex Mode · CSA2 Reindex Mode
Reuse Mode · Cross-layer Top-K index reuse, CSA2 Reuse Mode
Indexer Warmup · indexer warmup with full attention
Single-head index key · single-head key for index branch, single-head K_idx
Index Branch output · Index Branch value head output
Index Branch value head · index value head
fixed-pattern sparse attention 2
Local Block · forced sink and local window, forced sink and fixed local selection
linear attention & state space 16
Linear attention · linear attention mechanism
gated delta network 11
Gated DeltaNet · Gated Delta Net (GDN) module, Gated Delta Network linear attention
Kimi Delta Attention · KDA linear attention layers, Kimi Delta Attention (KDA)
KDA · Kimi Distributed Attention (KDA)
Hybrid linear attention · Hybrid Linear Attention Mechanism
Gated DeltaNet with bounded sigmoid output gate · Sigmoid output gate for Gated DeltaNet, bounded sigmoid gate
Kimi Delta Attention with input-dependent full-rank output gate · Input-dependent full-rank output gating, input-dependent full-rank projection
Kimi Delta Attention with lower-bounded log-decay · lower-bounded decay
hybrid layer stacking 13
Hybrid Attention · Hybrid attention mechanism, Hybrid Sparse and Linear Attention
Hybrid Mamba-Transformer · hybrid Mamba‑Transformer MoE, Mamba-Transformer
Mamba-2, Mixture-of-Experts, and Selective Attention Hybrid · Hybrid Mixture-of-Experts architecture, Interleaved Mamba-2 and MoE layers
Gated DeltaNet and Full Attention · Hybrid Attention Architecture
Gated DeltaNet and Qwen Sparse Attention · Hybrid Attention with QSA
Hybrid sparse mixture-of-experts Transformer architecture · hybrid sparse MoE Transformer backbone
Search-Based Sliding-Window Attention Pattern · SWA Pattern (Search-Based)
Dense Attention Fallback · dense attention as a per-layer fallback
channel mixer 68
Contextual gating for N-gram embedding injection · contextual gating mechanism
dense feed-forward network 6
SiTU-GLU · SiTU-GLU activation function, SiTU-GLU activation (soft-capped SwiGLU)
SwiGLU · gated SwiGLU activation, gated SwiGLU [9] activation function
SwiGLU clamping · SwiGLU activation with clamping, SwiGLU activation function with clamping
Dense feed-forward network · dense Feed-Forward Network (FFN)
SiTU · Sigmoid Tanh Unit (SiTU)
SwiGLU clipping · SwiGLU clipping for activation control, SwiGLU-clip
mixture of experts 61
Mixture of Experts · 198B total params / 11B activated params, 552B-parameter MoE
Sparse expert activation · Active parameters (sparse activation), active parameters
DeepSeekMoE · DeepSeekMoE framework
Mixture-of-Experts layers · MoE layers
MoE with routed and shared experts · 256 routed experts and 1 shared expert, Routed experts with shared expert
Asymmetric input/output activation split · asymmetric split
Dispatch recomputation · recomputing dispatch
Gated DeltaNet MoE · gated-delta-networks MoE architecture, routed and shared experts in MoE
Hybrid Mixture of Experts · Hybrid Attention Mixture-of-Experts (MoE)
MoE with 256 experts and top-8 routing · Sparse MoE with 256 experts and 8 top-k
Routed expert output modulation · routed expert modulation
expert routing 15
Top-8 expert routing · 192 experts, top-8 activated, top-8 routed experts
Token-level expert routing · token routing through 8 of 288 experts, routes each token through 8 of 288 experts
Routed experts · 192 routed experts (top-8)
Token-choice routing with softplus gating · Token-choice router with softplus gating
Top-6 expert routing · each token is routed to 6 of 256 experts
Fixed Top-k routing with frozen bias · fixed Top- k selection with a frozen bias
DP-aware routing · DP-aware routing mechanism
Token-choice routing · Token-choice mixture-of-experts routing
Top-4 expert routing · top-4 expert selection, top-4 routed expert selection
expert load balancing 15
Auxiliary-loss-free load balancing · auxiliary-loss-free strategy, Auxiliary-loss-free MoE load balancing
Quantile Balancing · QB load balancing for mixture-of-experts, QB for MoE load balancing
Modality-specific auxiliary-loss-free load balancing · modality-specific load balancing
Auxiliary-loss load balancing · auxiliary loss
formHC · formHC (MoE routing/load-balancing method)
Histogram-based quantile estimation · Histogram estimation, B uniform bins
Round-robin load balancing · --load-balance-method round_robin
MaxVio · MaxVio metric
fine-grained experts 2
DeepSeekMoE shared and fine-grained routed experts · shared and fine-grained routed experts
Granular Mixture of Experts · Granular MoEs
latent mixture of experts 5
LatentMoE · Latent Mixture of Experts, Latent Mixture-of-Experts (LatentMoE)
Normalized LatentMoE · Stable LatentMoE, Stable LatentMoE framework
positional encoding 14
Relative attention · Relative positional embedding
No Position Encoding · No Position Encoding on MLA layers, No Position Encoding (NoPE)
Omitting RoPE in attention layers · do not use RoPE in attention layers
Per-layer-type rotary position scales · per-layer-type rotary scales
YaRN · Modifying the model configuration file, Static YaRN context length extension
Rotary Position Embedding · RoPE, Rotary Positional Embedding (RoPE)
2D rotary position embedding · 2D rotary position embeddings (2D-RoPE), 2D-RoPE
Proportional Rotary Position Embedding · Proportional RoPE (p-RoPE), proportional rotary position embeddings
RoPE scaling · RoPE scaling for context extension, RoPE scaling techniques
Gated attention with partial RoPE · Gated attention with partial RoPE (50%), + GA Partial RoPE (50%)
normalization & residual 28
Attention Residuals (AttnRes) · Attention Residuals, AttnRes
Gated Residual · Gated Residual (GR), Gated residual connections
Identity Hyper-Connections · iHC (identity Hyper-Connections), Identity Hyper-Connections (iHC)
Hyper-Connections · Hyper-Connections (HC)
Per-Branch Scalar Write Gate · Per-branch scalar residual write gate, write gate
Single-Pass mHC · Single-PassmHC residual-stream mixing, Single-PassmHC
Block AttnRes · block attention residual, Block Attention Residuals
Elementwise Data-Dependent Residual Read Gate · read gate
Four-Branch Residual Stream · widens the residual stream into 4 branches
GatedNorm · GatedNorm elementwise self-gating, elementwise self-gate after RMSNorm
Independent Per-Branch Normalization · group RMSNorm over the widened stream
Two-Phase Block AttnRes Schedule · two-phase Block AttnRes kernel schedule, two-phase schedule
Widened Residual Stream · widening the residual stream
RMSNorm · RMSNorm normalization, root mean square normalization (RMSNorm)
Data-Dependent Residual Read and Write Operators · making and data-dependent
Pre-LN · Pre-LN (pre-normalization) placement, Pre-LN placement
Simplified AltUp · Simplified AltUp residual widening, a simplified variant of AltUp
QK-Clip · QK-Clip technique, QK clipping for activation control
Sparse Gated Residual Writes · sparse writes
prediction head 5
Multi-Token Prediction · 3.8B MTP layer, MTP
Multi-Token Prediction for residual codebooks · multi-token prediction (MTP) module
Nemotron Hybrid Multi-Token Prediction · nemotron_h_mtp
multimodal architecture 57
Encoder-free multimodal architecture · Encoder-free decoder-only architecture, encoder-free architecture
Early fusion multimodal training · Early fusion training on multimodal tokens, Unified Vision-Language Foundation
Discrete token encoding for audio · discrete token encoding, discrete token encoding for audio inputs
Hierarchical patch encoder · Hierarchical patch encoder for images
Mixed-modality training · mixed modalities from the start
Native multimodal understanding · native multimodal in one model, native omnimodal architecture
Multimodal mixture-of-experts Transformer · multimodal mixture of experts
Native multimodal encoding · natively multimodal
Native visual and audio understanding · multimodal understanding
Single shared multimodal backbone · single shared backbone
Vision encoder and multimodal aligner · vision encoder and aligner
Vision encoder–MLP projector pathway · vision encoder and an MLP projector
Visual modules for multimodal understanding · incorporating visual modules
Raw audio projection into the LLM embedding space · raw audio chunk projection
Vision Transformer (ViT) · Vision Transformer, ViT vision encoder
Audio Transformer (AuT) · attention-encoder-decoder model AuT
Data-parallel-first multimodal encoding · encoding runs data-parallel first
Decoupled Encoder Process (DEP) · decoupled encoder process
Disaggregated encoder training · disaggregated encoder design
Dual-format coordinate supervision · coordinate supervision
MoonViT-V2 · MoonViT-V2 vision encoder
Multimodal input · image input
Residual Vector Quantization (RVQ) speech representation · RVQ-based speech representation
Reusing the image token for video frames · reusing image token for video frames
Single-matrix-multiplication vision projection · single matmul vision projection, single large matmul projection for vision
Two-layer MLP vision projector · two-layer MLP projector
Two-stage text-then-multimodal pre-training · two-stage pre-training strategy
Universal Speech Model (USM)-based audio encoder · Universal Speech Model-based audio encoder, USM-based audio encoder
Variable aspect-ratio image handling · variable aspect ratios
context capacity 10
N-gram embedding lookup · Local-context N-gram embedding lookup, N-gram Embedding
1M-token context window · 1M context length, 1M context
N-gram embedding host-memory offload and prefetch · asynchronously offloaded to host memory, asynchronous prefetching
Engram · Engram conditional memory, Engram conditional memory module
Single-layer N-gram embedding placement · A single N-gram embedding layer
Progressive context extension · context window extension, progressive context extension curriculum
training objective 29
filed at the root 1
SigLIP sigmoid contrastive loss · sigmoid contrastive loss
language modelling objective 6
Cross-entropy objective · Cross-entropy training objective
Next-token prediction · next-token prediction objective
Fill-in-the-middle (FIM) completion · fill-in-the-middle completion, FIM Completion
multi-token prediction objective 6
Multi-Token Prediction · Multi-Token Prediction (MTP), Multiple-Token Prediction
Multi-Token Prediction Boosting · MTP Boosting, MTP-boosting phase
Multi-step MTP training · MTP: trained with multi-steps
Shared-weight Multi-Token Prediction · shared-weight MTP objective
distillation objective 7
Temperature-scaled forward KL distillation · temperature-scaled forward-KL loss
Dense-attention distillation · dense distillation
Quantization-aware distillation · Quantization-aware distillation (QAD)
Reverse KL divergence · reverse KL divergence loss
auxiliary loss 9
KL alignment loss · KL-divergence loss, KL divergence loss for indexer
Domain-specific KL regularization strength · Domain-adaptive KL regularization strength, varying strengths of KL regularization
Gradient detachment · Gradient Detach, gradient detach for KL loss
Language modeling loss plus KL loss · LM Loss + KL Loss
Loss masking · Loss masking of flagged trajectory content, mask
MoE sequence auxiliary loss · MoE sequence auxiliary loss coefficient
Unfinished-trajectory loss masking · mask the loss on unfinished trajectories
optimization 155
filed at the root 4
Updated hyperparameter scaling law · refit the scaling law, dedicated scaling-law studies
Phase-by-phase training cost accounting · phase-by-phase cost accounting
optimizer 29
Muon · Muon optimizer, Muon matrix-based optimizer
Category-specific assignment of Muon and AdamW · Tailored Training Recipe, division of labour between Muon and AdamW
Split fused gradients before orthogonalization · splitting of fused parameters
AdamW for the MoE router · we use AdamW for the router
Per-Head Muon · head-wise Muon, Per-Head Muon optimizer
Eight-step Newton–Schulz iteration · Eight-step Newton–Schulz orthogonalization, Newton–Schulz iteration to 8 steps
Sinkhorn-balanced update · Momentum-based Sinkhorn-balanced optimizer
AdamW · AdamW optimizer
Nesterov momentum · Nesterov-accelerated momentum, Nesterov trick
Asynchronous Micro-Group pipeline · an asynchronous Micro-Group pipeline
CUDA graph capture of the optimizer step · We capture the whole step in a CUDA graph
Muon orthogonalization accuracy refinement · orthogonalization accuracy
Peer-to-peer shard retrieval for Muon orthogonalization · P2P communication
Selected number of Newton–Schulz iterations · the number of iteration steps
learning-rate schedule 13
Cosine decay · Cosine decay learning rate schedule, cosine learning rate schedule
Warmup-Stable-Decay · Warmup Stable Decay learning rate schedule, Warmup Stable Decay (WSD)
Batch-size warmup · Ramping the batch size over early training, Batch-size warmup during early training
Engram learning-rate scaling · learning rate of Engram is scaled by 5×
Linear warmup followed by cosine decay · linear warmup then cosine decay
Scaling-law fit · new scaling-law fit
Scheduled batch-size growth · batch size scheduling strategy
WSD-specific scaling law · WSD learning-rate scaling law
training precision 19
NVFP4 · NVFP4 4-bit Training Format, ultraefficient 4-bit NVFP4 training format
NVFP4 pre-training · NVFP4 pre-training recipe, Pre-training with NVFP4 quantization
BF16 · BF16 precision, BF16 testing
FP8 mixed-precision training · FP8 mixed precision
FP4+FP8 mixed precision · FP4 + FP8 Mixed
FP8-precision reinforcement learning · RL was done in FP8 precision
MXFP8 · keep these layers in MXFP8
E2M1 · E2M1 Floating Point, E2M1 datatype
FP8 storage for the residual state · FP8 storage for the widened residual state, keep the residual state in FP8
Mixed-FP8 quantization · mixed-FP8
Mixed-precision training · mixed-precision training/rollouts
quantization-aware training 15
Quantization-Aware Training · Quantization-Aware Training (QAT)
Quantize-Dequantize Training · quantize–dequantize (QDQ)
FP4 Quantization · FP4 (MXFP4) quantization, MXFP4 quantization
FP4 Quantization-Aware Training · MXFP4 quantization-aware training, MXFP4 quantization-aware training (QAT)
MXFP4 Weights with MXFP8 Activations · MXFP4/MXFP8 quantization-aware training
NVFP4 Training · quantized pretraining with NVFP4
Stochastic Rounding for Mamba Cache · mamba-cache-stochastic-rounding
Stochastic Rounding of Gradients · stochastic rounding on gradients
training stability 17
Cross-replica model-weight hash consistency checks · Cross-replica Hash Checks
Exponential moving average of checkpoints · exponential moving average (EMA)
Learning-rate elevation for training stability stress testing · raising the learning rate
Off-policy sample filtering · Dropping off-policy and noisy samples
Training stability stress testing · stress tests
Weight clipping · weight-clipping mechanism
training parallelism 29
MoonEP · MoonEP expert placement planning
Sequence Parallelism for Activations · sequence parallelism (SP) for activations
Expert Parallelism · MoE Expert Parallelism (EP), expert parallel
Dynamic Context Parallelism for Large Multimodal Samples · Dynamic CP in multimodal encoder
GPU Planning Kernel for Redundant Expert Migration · online planning of redundant experts
KDA Context Parallelism · KDA Context Parallelism (KCP)
Load-Balanced Image Sharding · Balanced image sharding
NS-FLOP-Balanced Static Parameter Partitioning · An -balanced static partitioner
pipeline-bubble scheduling of ViT computation · scheduled into pipeline bubbles
SConv-Aware Tensor-Parallel Sharding · sconv-aware TP sharding
SM-Level Context Parallelism · intra-device context parallelism
training runtime 29
Asynchronous checkpointing · Asynchronous background checkpointing
Asymmetric local-read and single-writer cache paths · Asymmetric read/write cache paths
Autotune configuration generation · autotune
Batch-level embedding prefetch · Embedding prefetch
Caching the distributed checkpoint save plan · cache the distributed save plan
CPU-resident optimizer states · Optimizer states stay in CPU memory
Cross-rank remote activation offloading · remotely offload activations
GPU-to-GPU weight synchronization over GPUDirect RDMA · GPU-to-GPU weight transfer
Job-level eviction and reclaim · Per-job eviction and reclaim
Overlapping NCCL transfers with device-to-host copies · overlapped NCCL transfers with D2H copies
Persistent shared-storage cache for compiled artifacts · Persistent warm cache on shared storage
Post-iteration NVMe offloading of training states · offload training states ... to NVMe
Pre-admission hardware stress testing · pre-flight checks
Same-node sticky pod respawn · Sticky pod respawn
Seeding node-local storage from a warm shared cache · Node-local seeding at startup
data curation 160
filed at the root 5
Training data deduplication and filtering · deduplication and filtering
Multilingual corpus expansion · expands its multilingual corpus
data sourcing 17
GitHub Crawl · GitHub Crawl via REST and S3 APIs
High-recall web-data curation · high-recall web data
Streaming training-data ingestion · stream training data
Web knowledge graph construction and question generation · web knowledge graph construction
Video sampling parameters · mm_processor_kwargs
data filtering 37
Conservative model-based noise filtering · model-based and deliberately conservative
Quality-bucket sampling · samples from corresponding quality buckets
CSAM filtering · child sexual abuse material filtering
CBRN pre-training data filtering · CBRN safety filtering of pre-training data, CBRN safety pre-training data filtering
Filtering batched auto-generated and templated content · Quality filtering against model collapse
Heuristic filtering · Heuristic filtering for multilingual data
Keyword- and regex-based filtering · Keyword- and regex-based content filtering, targeted keyword- and regex-based filters
Pathological repetition filtering · n-gram repetition filtering
Data cleaning · fine-grained data cleaning
Dense annotation of ambiguous low-quality data · densely annotate this region
Image-text relevance filtering · image-text relevance threshold
PCA-based dimension selection for decorrelated quality signals · PCA-informed subset of its dimensions
Post-training data filtering for safety and factuality · post-training data filtering
Pre-training data filtering for decontamination and safety · data filtering
SmolVLM-based image-text quality scoring · SmolVLM
Structural checks for malformed examples · structural checks
Two-axis document-quality labeling by noise and information · two complementary axes
deduplication 5
In-pack deduplication · in-pack deduplication constraint
Semantic image deduplication · deduplicate them based on image semantics
synthetic data 56
Automatic synthesis of task-oriented RL environments · automatic environment-synthesis agent, Large-Scale Agentic Tasks
Modular synthetic-data pipeline composition · modular approach to synthesis
Knowledge distillation for synthetic data · knowledge distillation, distillation
Corpus rephrasing with fidelity verification · rephrasing recipe
End-to-end synthetic environment generation · Synthetic environment generation pipelines
Heterogeneous answer-generation agents · answer-generation agents
Metadata-conditioned synthetic generation · Help LLMs through metadata
Multi-agent task-attempt validation · multiple distinct agents attempt the task
Multi-stage synthetic-data generation cascade · Multi-stage cascade
Programmatic multimodal data generation · programmatic multimodal data
Search-based question construction · question-construction agent
Search-capable answer verification agent · verification agent
Synthetic data generation and augmentation · synthetically generated or augmented
Synthetic system-message augmentation · synthetically generated system messages
Targeted synthetic data rephrasing · targeted synthetic rephrasing at scale
Task-specific data injection during mid-training · Task-specific data injection
Two-sided test validation for synthetic coding tasks · two-sided correctness check
data mixture & curriculum 26
AutoMixer · Automated data-mixture optimization
Vulnerability-discovery data inclusion · added vulnerability-discovery data
Agent-centric multimodal data mixture · agent-centric data mixture
Deterministic pre-splitting of ultra-long documents · deterministically pre-split
Domain-specific multimodal datasets · domain-specific datasets
Dynamic sampling for mixed-task reinforcement learning · Dynamic sampling for mixed-task RL, dynamic sampler
Gaussian-based data mixture and curriculum construction · Gaussian-based approach
Interleaved multimodal training · interleaved training
Long-context data upsampling · upsampling long-context data
Pass-rate-based RL task sampling · pass-rate buckets
Smooth weighted round-robin source scheduling · smooth weighted round-robin
Three-stage pre-training curriculum · three distinct stages
Union-based integration of text-only and multimodal corpora · union of both data sources
ProtocolQA-aligned RL training datasets · Aligned RL datasets with ProtocolQA
sequence packing 6
Length-aware best-fit packing · length-aware best-fit packing strategy
Best-fit packing · best-fit packing algorithm, best-fit sequence packing
Interleaved image-text sequences · interleaved image-text data construction
tokenization 8
o200k_harmony tokenizer · o200k_harmony Byte Pair Encoding tokenizer
Quick Instruction tokens · Quick Instruction
Separate PT and IT end tokens · PT versus IT formatting
post-training 267
filed at the root 10
Scalable reinforcement learning post-training · Scalable Reinforcement Learning Framework
Reinforcement learning for low-pass-rate tasks · RL reserved for tasks with low pass rate
Three-stage post-training with multi-teacher on-policy distillation · three-stage post-training paradigm
Three-stage Thinker post-training strategy · three-stage strategy for the Thinker
supervised fine-tuning 24
LoRA fine-tuning · LoRA adapter fine-tuning, LoRA adapter
Light SFT on teacher-distribution data · teacher-distribution SFT warmup
SFT checkpoint for RL research · SFT starting point for RL research
Supervised fine-tuning · SFT, Supervised Fine Tuning (SFT)
Rejection sampling · rejection sampling fine-tuning
Self-correction cold start · Cold start from self-correction
Evaluation-based early stopping · evaluation-based early stopping during SFT, early stopping based on evaluation scores
Prompt-based cold start for tool use · Cold-Start
Safety training to an internal specification · Safety training to internal spec
Specialist-model training · Specialist training, independent specialist models
Token-budget-based SFT data blending · token-level target proportions
Fine-tuning · fine-tune
reinforcement learning algorithm 62
Group Relative Policy Optimization · GRPO, GRPO (Group Relative Policy Optimization)
Groupwise Advantage Redistribution · Groupwise Advantage Redistribution (GAR)
Chain-of-Thought Reinforcement Learning · similar CoT RL techniques as OpenAI o3
Mixed Reinforcement Learning · Mixed RL Training
Reinforcement Learning · RL, Reinforcement Learning (RL)
IcePop · IcePop technique, token-level masking using IcePop strategy
Off-Policy Sequence Masking · Off-policy sequence masking for GRPO
Reinforcement Learning from Verifiable Rewards · unified RLVR, RLVR
Reinforcement Learning Post-Training · RL post-training, larger-scale RL post-training
Rollout Routing Replay · Rollout Routing Replay (R3)
Abstention Training · abstention-focused reinforcement learning
Advantage Shaping · Advantage shaping on flagged tokens
Domain-Specialized RL Experts · scaling RL across three broad domains
Flagged-Token Masking and Penalty · flagged tokens
Freezing the MoE Router During Reinforcement Learning · freeze the router for RL training, freeze the MoE router
Group-Wise Policy Optimization · group-wise policy optimization algorithm
Hierarchical Penalty Escalation · Penalties escalate along the hierarchy
Interaction-Aligned Reinforcement Learning · Interaction-Aligned RL
Keep Candidate-Set Replay · candidate-set replay
Large-Scale Reinforcement Learning · Reinforcement learning for post-training
Per-Token Tool-Error Step Penalty · Tool-error step penalty
Pivot Reinforcement Learning · PivotRL
PPO-Style Policy-Ratio Clipping · PPO-style clipping
Quality-Weighted Advantage Redistribution · quality factors
Reinforcement Learning from Code Execution Feedback · RLCEF tasks
Reinforcement Learning Training · RL training
SAO with Compaction · SAO with compaction (RL strategy)
Single-Harness Reinforcement Learning · single-harness RL
Moonlight Scaling · Moonlight scaling in RL
reward modelling 35
Groupwise Reward Synthesis · Groupwise Reward Synthesis (GRS)
Binary terminal-verifier reward · binary verifier
Generative Reward Model · GenRM, Generative Reward Model (GRM)
Length-adjusted RL reward · length penalty reward, length penalty
Verifier cross-checking · verifier cross-checks
Abstention-aware reward for factual QA · Abstention-aware rewards for factual QA
Agentic Generative Reward Model · Agentic Generative Reward Model (GRM)
Collaboration bonus · Collaboration bonus in RL reward
Hack-agent screening · Hack Agent
Monitoring-only penalty strategy · monitor
Outcome Reward Model · Outcome Reward Models, Outcome Reward Models (ORMs)
Principle-conditioned Generative Reward Model · principle-following GenRM
Rule-based outcome reward · rule-based outcome rewards
Rule-based verifier · rule-based verifiers
Training-time trajectory auditing · Training-Time Auditing
preference optimization 4
Direct Preference Optimization · Direct Preference Optimization (DPO)
policy distillation 19
Multi-Teacher On-Policy Distillation · Mixture of On-Policy Distillation, MOPD
On-Policy Distillation · on-policy distillation (OPD)
Specialist Distillation · multi-domain specialist distillation, specialist-model distillation
Prefix-Conditioned On-Policy Distillation · Prefix-Conditioned OPD, SFT-Prefix OPD
Asynchronous Multi-Teacher On-Policy Distillation · asynchronous MOPD
Autonomous Student Rollouts · Standard MOPD
Distillation Fine-Tuning on MiMo-Generated Data · trained on MiMo-generated data
IcePop Token-Level Loss Masking · token-level masking using IcePop strategy
Model Distillation · distillation
Per-Token On-Policy Distillation Reward · per-token OPD reward
rollout & RL infrastructure 76
Asynchronous reinforcement learning · asynchronous RL architecture, asynchronous training paradigm
Agent-centric rollout execution · agent-centric execution model
Asynchronous reinforcement learning framework · asynchronous RL framework
One-step off-policy asynchronous reinforcement learning · one-step off-policy asynchronous RL setup
Partial rollout · Partial rollout for mixed-task RL, partial rollout for synchronous RL
Asynchronous reinforcement learning infrastructure · asynchronous RL infrastructure
SLIME · SLIME (Reinforcement Learning framework)
Token-in-token-out (TITO) · token-in-token-out, token-in, token-out (TITO) API design
Asynchronous sample generation · asynchronous generation of samples
Blocking in-flight rollout steps on weight updates · blocking in-flight rollout steps on update
Bounding the maximum off-policy ratio · bound the maximum off-policy ratio
Closed-loop multi-turn rollout · Multi-turn rollout
Composite early-stop strategy · early stop strategy
Decoupled control plane and data plane · Disaggregated Data Plane and Control Plane
Dialogue prefix matching · Prefix matching
Dynamic training recipe reconfiguration · dynamic reconfiguration during training
Long-horizon rollout budgets · long-horizon rollout budgets in RL
Multi-harness rollouts · multi-harness rollout in RL
Multimodal rollout data handling · Multi-Modal Data
Per-source oversampling allocation · oversampling ratio
Per-source priors for rollout sequence-length estimation · per-source priors
Persistent host-actor pools · fixed-size pools of persistent host actors
Sample Mixer · A sample mixing mechanism
Token-level interruption · Token-level interruption of generation
Batch-level rollout dispatch · batch-level dispatch of rollout prompts
Prompt-level dispatch · prompt-level dispatch of rollout prompts
agentic post-training 31
Agentic reinforcement learning · agentic RL, Large-scale agentic reinforcement learning
Agentic tool-use training · agentic tool-use post-training
Environment preparation to prevent solution leakage · Environment Preparation
Multi-environment reinforcement learning from verifiable rewards · multi-environment RLVR
Multi-harness training · Multi-harness reinforcement learning
Python tool use in chain-of-thought · python tool
Reinforcement learning on synthetic agentic data · large-scale RL on synthetic data
Training agentic policies with real-world tools · real-world tools
Unified trajectory representation for agentic reinforcement learning · unified trajectory representation
Unified white-box reinforcement-learning environment · unified white-box RL environment
Reinforcement learning restricted to search and code environments · RL only in search and code environments
mid-training & continual pretraining 6
Continued training for visual understanding · continued training
Continual pretraining · continued pre-training, Continuous Pretraining
Continual pretraining for long-context extension · long-context extension, long-context training
inference & serving 375
filed at the root 1
decoding strategy 28
Speculative decoding · speculative decoding methods, speculative decoding module
Multi-Token Prediction · MTP-based speculative decoding, Multi-Token Prediction (MTP)
DSpark · DSpark speculative decoding, DSpark speculative-decoding module
DFlash · Block-6 DFlash speculative decoding, block-6 DFlash for speculative decoding
Fused recurrent replay kernel · single fused kernel
EAGLE · EAGLE (speculative decoding), EAGLE (speculative decoding algorithm)
NEXTN speculative decoding · native NEXTN speculative decoding, NEXTN speculative decoding algorithm
Presence Penalty · Presence penalty adjustment, Presence-penalty tuning
Best-of-N scaffolding · Best of K scaffolding, Best-of-n selection
Multi-stage candidate filtering · multi-stage filtering pipeline
EAGLE-3-style draft-model fine-tuning · Draft Model Fine-Tuning
Longest-trace selection · longest thinking trace selection
Top-k over token clusters · top-k operation on clusters of tokens
Multi-layer EAGLE · enable-multi-layer-eagle
Task-specific sampling parameters · sampling parameters
reasoning control 42
Configurable reasoning effort · dial 'thinking effort' up or down, reasoning_effort request field
Chain-of-thought reasoning · Chain-of-Thought, COT
Default thinking mode · Default thinking mode for generation, thinking mode by default
Interleaved thinking between tool calls · Interleaved thinking, per-request control via enable_thinking
Always-on thinking mode · always-on thinking block, always has thinking enabled
Control-token-enabled thinking mode · configurable thinking modes
Capped linear reasoning-token length deduction · length deduction
Effort-dependent exponential token-penalty schedule · effort-dependent token-penalty coefficient, exponential token penalty schedule
Thinking mode selection · Thinking mode, thinking
Maximum thinking effort · Max thinking effort, max thinking
clear_thinking chat-template parameter · clear_thinking chat template flag, clear_thinking
Deployment-time scalar effort control · Scalar effort control of response length, scalar efforts
Cross-turn persistent reasoning history · persistent reasoning-history chat template, persistent thinking history
Inference-time reasoning budget control · inference-time budget control
Reasoning parser · reasoning parser step3p5
Generate-verify-refine loop · Generate-verify-refine test-time scaling, generate–verify–refine methodology
Test-time compute scaling · test-time scaling, test-time-scaling framework
Per-problem reasoning-budget control · per-problem budget control mechanism
Thinking modes (off and max) · two modes: off and max
Turn budget capping · turn budgets
Disabling reasoning via chat-template configuration · disabling reasoning mode via chat template, Disable Reasoning
Configurable thinking or reasoning mode · configurable thinking/reasoning mode
reasoning-enabled inference toggle · reasoning enabled boolean
Task-appropriate maximum output length · Adequate Output Length
Parallel-fewest-step sampling · parallel scaling, Parallel-fewest-step
Parallel test-time compute scaling · in parallel
Reward-optimized preferred reasoning length · preferred reasoning length
Serial test-time compute scaling through context management · serially through context management
KV cache management 42
Compressed KV caching · KV cache compression, Smaller KV cache
Block-based KV cache · block-based key-value cache
Cross-layer KV-cache reuse · Cross-layer key-value cache reuse
Customized heterogeneous KV-cache layout · customized KV cache layout
Fine-grained prefix hashing · Prefix hashing runs on fine hash blocks
KDA-aware prefix-cache management · joint KDA–MLA prefix cache management, KDA-aware prefix cache
Persistent per-dialogue-context KV caching · Context Caching
Projected-input caching for speculative KDA rollback · cache only these projected inputs
Write-back external KV-cache policy · write-back design
Chunked prefill · --chunked-prefill-size 16384, chunked prefilling
Automatic cache · Full automatic Cache support
Prefix caching · prefix cache
Language-model-only serving mode · language-model-only, Text-Only
RadixCache · radix cache
Disable prefix caching for benchmarking · disabling prefix caching
Hierarchical KV caching · hierarchical extension of the cache
Inference-side KV-cache reset on weight synchronization · KV-cache reset on weight sync
Key-value reuse in global attention layers · reuse of keys as values in global layers
Least-recently-used eviction · LRU eviction policy
On-disk KV-cache storage · on-disk KV cache storage mechanism
Overlapped host-memory prefetching · host-memory prefetching
Stateless in-memory prefix caching · stateless prefix caching in memory
Periodic cache checkpointing · cache checkpointing, Periodic checkpointing of SWA KV entries
Mamba prefix caching in align mode · prefix caching in align mode for Mamba
inference quantization 64
Quantization · quantization algorithms, quantized versions
FP4 KV-cache quantization · FP4 main KV cache, FP4 KV caching
Mixed-FP8 layers in an NVFP4 recipe · mixed-FP8 layers
Mobile-specialized quantization schema · custom mobile-quantization schema
NVFP4 quantization for routed-expert GEMMs · NVFP4 routed-expert GEMMs
Offline weight-layout permutation · weight layout is permuted offline
FP8 · FP8 (8-bit floating point), FP8 quantization
NVFP4 quantization · NVFP4 4-bit floating-point quantization, NVFP4 checkpoint
NVFP4 KV-cache quantization · NVFP4 KV cache, NVFP4 KV caching
Post-RoPE KV-cache quantization · quantizing the cache after RoPE
FP8 KV-cache quantization · FP8 KV cache during RL rollouts, FP8 E4M3 key-value cache
MXFP4 weight quantization · Microscaling FP4 (MXFP4) weights, MXFP4 weights
Post-training quantization · Post-Training Quantization (PTQ)
Four-Over-Six · Four-Over-Six FP4 weight scale selection, Four-Over-Six scale selection method
GGUF · GGUF (GPT-Generated Unified Format), GGUF quantization
MXFP8 activation quantization · Microscaling FP8 (MXFP8) activations, MXFP8 activations
SSM cache quantization · quantize the Mamba SSM cache
AWQ INT4 weight quantization (W4A16) · INT4 AWQ weight quantization
Block-wise E4M3 FP8 weight quantization · FP8 quantization with block-wise e4m3, native FP8 (block-wise e4m3) weights
Embedding and KV-cache quantization · Embedding and KV cache optimization
FP16 cache storage with stochastic rounding · FP16 Cache with Stochastic Rounding
FP4 precision for the attention indexer · FP4 precision for indexer, FP4 precision for attention indexer
FP8 inference · FP8 rollouts
FP8 low-precision speculative draft computation · Low-precision draft computation (FP8)
Heuristic mixed per-layer precision quantization · heuristic mixed per-layer precision recipe
Mixed-precision INT4/INT8 layer-wise quantization · mixed-precision quantization strategy
MXFP4 tensor packing · pack every two values in one uint8 value
NVFP4 quantization with modelopt · NVFP4 model
NVIDIA ModelOpt quantization · modelopt quantization
Post-quantization of MoE model layers · quantization of MoE layers
Random Hadamard transform · Random Hadamard Transforms
Static activation quantization · Static activations
W4A16 quantization · W4A16 (weights NVFP4, activations BF16), W4A16
W4A8 quantization · W4A8 mixed-precision quantization strategy
NVFP4 ModelOpt re-quantization · NVFP4 ModelOpt re-quantization (W4A16), ModelOpt re-quantization
Q4_0 quantization · Q4_0 blockwise quantization
Block-scaled INT8 quantization with stochastic rounding · Block-scaled INT8 quantization
Max-based scaling · Max-based quantization scale calibration, Max-based weight scaling
MSE-based scaling · MSE-based weight scaling
Empirical bits-per-element budget selection · Bits per Element Selection
Low-bit quantization · low-bit methods
In-flight block-wise FP8 weight quantization · in-flight block-wise weight quantization
serving parallelism 17
Zero-copy fused token permutation and unpermutation · zero-copy communication, fused per-mute/unpermute operator
Tensor parallelism (degree 4) · Four-way tensor parallelism for serving, tensor-parallel-size 4
DeepEP · --moe-a2a-backend deepep, DeepEP expert parallelism across nodes
Prefill-decode disaggregation · Prefill–Decode (PD) disaggregation
Attention data parallelism · Attention Data Parallelism (DP), Data Parallel attention
Tensor parallelism (degree 8) · tensor-parallel serving on 8 GPUs, tensor parallel on 8 GPUs
Data-parallel vision encoding · --mm-encoder-tp-mode data
Tensor parallelism for MoE layers · tensor parallelism in MoE
TensorRT-LLM all-reduce backend · trtllm all-reduce backend
Expert parallelism · enable expert parallelism
inference scheduling 20
Prefix-cache-aware session affinity scheduling · cache-aware affinity scheduling
Request-class resource-budget admission control · budget-based admission control
Runtime-signal-based rollout concurrency auto-throttling · auto-throttling mechanism
Workload-aware routed-expert GEMM scheduling · workload-aware scheduler
Asynchronous scheduling · async scheduling
NUMA binding of workers to GPU-local CPU sockets · NUMA Binding for CPU-GPU affinity, NUMA Binding
Greedy feasible-rank placement by remaining capacity · greedy heuristic
Latency-sensitive execution class with priority isolation · a latency-sensitive (LS) execution class
Max sequences tuning · Max sequences tuning for context fitting, --max-num-seqs 32
Registering NVLink domains as Ray custom resources · registers it as a Ray custom resource
Continuous batching · Continuous batching for inference serving
Exacto routing · Exacto (highest tool-calling accuracy)
Increasing batch size for inference · increasing batch size for MoE inference
inference kernel 50
Batch-invariant deterministic kernels · Deterministic and batch-invariant kernels
Kernel fusion · inference kernel fusion, operator fusion
GPU planning kernel for online expert placement · GPU planning kernel
Separate-stream shared-expert GEMM overlap · separate stream
Shape-aware kernel dispatch for Hyper-Connection · Shape-aware dispatch
Synchronization-free static-shape MoE execution · sync-free MoE execution with static shapes, sync-free execution with static shapes
WarpDecode token-centric routed-expert decoding kernel · token-centric design of WarpDecode
CUDA Graph · CUDA Graph integration, the step is captured in a CUDA graph
CUDA Graph capture size reduction · reduce --max-cudagraph-capture-size
FlashAttention 3 · --attention-backend fa3, FlashAttention 3 (fa3)
TensorRT-LLM multi-head attention backend · trtllm_mha, trtllm_mha attention backend
DeepGEMM-based batch-invariant matrix multiplication · replace it end-to-end with DeepGEMM
FlashKDA · FlashKDA chunkwise kernel for KDA
Fused mHC kernels · fused kernels of mHC
Fused QSA kernel · Fused sparse-attention and KL-loss kernel
KDA algorithm–system co-design · systems co-design for KDA
Marlin NVFP4 kernels · Marlin NVFP4 linear and MoE kernels
Optimized Triton MoE kernel with MXFP4 support · optimized triton MoE kernel
Persistent kernel · persistent kernel rewriting
ROCm AITER sparse MLA attention backend · --attention-backend ROCM_AITER_MLA_SPARSE
Split-K CuTe GEMM for low-batch Hyper-Connection Mix · low-latency split-K CuTe GEMM
FlashAttention 4 · fa4
HPC-Ops attention backend · HPC_ATTN
FP8 GEMM · FP8 matrix multiplication (GEMM)
Avoiding split-K · we abandon split-k in most scenarios
context management 18
Preserved thinking history mode · preserved thinking, Thinking Preservation
Excluding prior thinking from conversation history · No Thinking Content in History
Discard-all context management · Discard-all, discard-all context management strategy
Context compaction · context compaction at 300K tokens, Context-compaction strategy
Context folding · context-folding strategy, simple context-folding strategy(256k)
Thinking context management for tool use · Thinking Context Management
Hierarchical context management · Hierarchical Context Management strategy
Keep-recent-k · keep-recent-k context management, Keep-recent-k strategy
Memory compression · memory compression (context consolidation)
Preserve thinking · Preserve thinking blocks
Modality-specific deployment · deploying only the modalities you need
Summary-based context compression · summary-based compression
agentic scaffolding 93
Function calling · Function calling (structured tool use), native function calling
Harmony format · harmony response format
Role-based instruction hierarchy · role hierarchy
Assistant output channels · channels
Sandbox snapshots · Snapshot
Sandbox state forking · Fork
Automatic tool choice · auto-tool choice, enable-auto-tool-choice
Browser tool · web browsing tool training
Generic tool-call parser · Tool call parser
XML-based tool-call schema with DSML token · XML-based tool-call schema
Asynchronous teammate spawning · spawn_teammate
Bash commands for context retrieval · Bash commands
Browsing tool with domain filtering · browsing tool with a domain block
Commentary-channel preambles · preambles on the commentary channel
Concurrent subagent orchestration · using more than 20 concurrent subagents
Current-turn-only reasoning-mode detection · scans only the current turn’s content
Dynamic tool loading · dynamically loaded tools
Fresh and fork teammate initialization modes · fresh mode or fork mode
Hy v4 tool-call parser · Tool-call parser hy_v4
Indexed parallel tool calls · tool and index attributes
Interleaving tool calls with chain-of-thought · interleaving tool calls within the CoT
JSON Schema response formats · structured output response formats
Jupyter Notebook code interpreter · Jupyter Notebook as a code interpreter
Lead-agent interruption of teammates · Lead-agent interruption of teammate turns, interrupt_agent
Prompt-based cold start for reasoning in tool use · Cold-Start
Prompt-enforced tool-call format · our designed toolcall format
Qwen3 XML tool-call parser · tool-call parser qwen3_xml
ReAct Toolbelt · ReAct Toolbelt framework
Request timeouts and retries for agent loops · set a request timeout and a retry
Sandboxed Python execution loop · secure sandboxed Python execution loop
Scrollable browser text window · scrollable window of text
Single-Bash-tool scaffold interface · singlebash tool
Stateful Python tool · stateful tool
Stateless Python tool reference implementation · stateless mode
Step3p5 tool-call parser · tool-call parser step3p5
Streaming-delta handling across block boundaries · Deltas spanning block boundaries
Thinking with tools · Thinking with tools capability
Tool output message format · tool message format
Tools section in the system message · # Tools section
Typed tool arguments · Arguments are typed
TypeScript-like function schema syntax · TypeScript-like type syntax for functions
XTML chat template · XTML-based chat template
Advisor strategy · Advisor Mode
MCP tool configuration · MCP configuration file for available tools, Tool use via MCP configuration
Checkpointing · Checkpointing for rollback of agent edits, checkpointing feature
Claude Code skills for Tinker · Claude Code skills
Compact TXT image-path notation · compact TXT image-path prompt encoding, compact <image>path</image> TXT notation
Qwen3 Coder tool-call parser · --tool-call-parser qwen3_coder
software implementation 79
filed at the root 7
Separate encoding and inference modules · separate encoding/ and inference/
fastsafetensors · fastsafetensors checkpoint loading, fastsafetensors load format
inference engine 15
vLLM · Serving MiMo-V2.6-Pro-RL with vLLM, vLLM model serving
TensorRT-LLM · TRTLLM, TensorRT-LLM inference optimization
Atlas inference library · Atlas inference library on vLLM
Environment variable propagation to subprocesses · Environment propagation
llama.cpp · llama.cpp inference library
vLLM or SGLang serving · vLLM/SGLang serving, vLLM or SGLang
vLLM tensor parallelism · vLLM at --tensor-parallel-size 8
SGLang · Serving MiMo-V2.6-Pro-RL with SGLang
vLLM language-model-only mode · --language-model-only
training framework 6
DAG-based asset lineage tracking · directed acyclic graph (DAG) of assets
Hugging Face Transformers · Transformers, HuggingFace Transformers
NeMo · NeMo 26.04.01
kernel & quantization library 18
FlashInfer · FlashInfer backend
AngelSlim · AngelSlim compression toolkit, AngelSlim toolkit
FlashInfer NVLinkOneSided · FlashInfer NVLinkOneSided All-to-All, FlashInfer's NVLinkOneSided
FlashInfer TensorRT-LLM backend · flashinfer_trtllm
NVIDIA Model-Optimizer · Model-Optimizer
AITER · AITER linear and MoE backends, AITER backend
Container-baked precompiled artifacts · Container-baked artifacts
FLASHMLA_SPARSE · FLASHMLA_SPARSE attention backend
Marlin MoE backend · moe-backend marlin
agent product 11
infrastructure service 22
Automatic provider failover · automatic retry on next-best provider, provider failover routing
Sandbox pause and resume · pause-and-resume sandbox lifecycle control, Pause and Resume
Durable peer mailbox · Durable peer mailbox communication
In-house batch scheduler · custom cluster scheduler
Local container-image caching · Container caching
NeMo Guardrails · NeMo Guardrails policy enforcement
NeMo Switchyard · NeMo Switchyard model routing
Per-node pooling of initialization actors · pooling initialization actors per node
Replacing short-lived Ray actors with tasks · converting short-lived actors to tasks
evaluation 109
filed at the root 35
Prompt-based output standardization · Standardize Output Format, standardize model outputs
Evaluation without safety filters · safety evaluation without safety filters, safety evaluations without safety filters
Refusal-suppressed variants for capability estimation · refusal-suppressed variants
avg@k · average over 3 rollouts (avg@3), avg@3
Almost@1 · Almost@1 rollout success metric
Behavior testing with stubs and edge cases · behaviour test for every bash script
Deadline-bounded rollout evaluation · explicit per-rollout wall-clock deadlines
Documenting evaluation configuration · document our exact settings
End-to-end exploit development evaluation · Exploit development (Tier 2)
Evaluation equivalence threshold · Evaluation equivalence threshold of 0.3
Fixed MTP draft-token acceptance length · fixed acceptance length
Four-run mean pass@1 · Mean pass@1 across four runs
Joint validation protocol · joint protocol
Mean@5 · Mean@5 evaluation metric
Prompt engineering to elicit answers · enforce an answer via prompt engineering
ProtocolQA robustness validation · ProtocolQA robustness checks
Refusal behavior quantification · Quantified refusal behavior
Standard function-calling format · standard function call format
Union-of-capabilities coverage scoring · union of capabilities across revisions
Visual checking of generated renders · visual checks
Vulnerability discovery with proof-of-concept evaluation · Vulnerability discovery (Tier 1)
V-model iterative testing and validation · V-model methodology
benchmark 25
Capture-the-Flag (CTF) evaluation · CTF challenges
Disabling tool search · disabling tool search in evaluation, Tool Search disabled
Domain whitelist · domain whitelist for agent environments
Fail-to-pass and pass-to-pass evaluation points · fail-to-pass and pass-to-pass points
GPU kernel optimization task suite · kernel optimization tasks
High-confidence ProgramBench subset · High-confidence benchmark subset filtering, high-confidence subset of ProgramBench
Multiple-rollout evaluation · Multiple-rollout evaluation per task, up to three rollouts per task
Verified CUDA kernels · Verified CUDA kernels for synthetic data
Computer use closed-loop evaluation · Computer use closed loop
evaluation harness 14
DeepSeek Harness Minimal mode · DeepSeek Harness Minimal evaluation mode, Minimal mode of DeepSeek Harness
Benchmark evaluation framework · benchmark framework
Cross-scaffold evaluation · Cross-scaffold robustness evaluation
Isolated container per rollout · isolated container per agent rollout
Pool · agentic coding evaluations run using pool
judge 16
Agent-as-a-Judge · Agent-as-a-Judge evaluation method
Decompositional instruction-following judge · dedicated IF judge
Deterministic tool-call verifier · deterministic verifier
GPT-5.5 (medium) judge model · GPT-5.5 (medium) as the judge model
LLM-based anti-cheating judgment · LLM-based judgement for anti-cheating, LLM-based judgement
LLM-based correctness judging · LLM-based judging pipeline
Mandatory agentic judge protocol · mandatory protocol for agentic judge
Official task verifier scoring · official separate verifier
Reward-hack detection judge · Reward hack judge
Rule-based anti-cheat checks · rule-based judgement
human & real-world evaluation 19
Held-out evaluation · held-out generalization gates
Blind side-by-side expert evaluation · blind expert judging of model outputs, blind expert judging
Multi-turn open-ended external red-teaming · Multi-turn open-ended red-teaming
Human evaluation · human evaluations
Internal engineering task evaluation · Internal engineering tasks evaluation
Pairwise comparison for Chinese writing · Pairwise comparisons for Chinese writing
Reward hacking mitigation in evaluations · reward hacking mitigation
Same-prompt comparative DevOps testing · same-prompt DevOps test
White-collar enterprise task evaluation · White-collar enterprise tasks evaluation
A/B testing on real tasks · A/B testing on three real tasks, A/B on three real tasks
other 16
filed at the root 16
Early Release with Feedback-Driven Iteration · ship early and hear what breaks
Prompt Modality Ordering · Modality order
Removing Git-History Leaks from Benchmark Images · removed to prevent possible reward hacking
Request Caching in the Browser Tool · the tool caches requests
Reward-Hacking Detection System · hacking-detection system
Score-versus-Cost Efficiency Comparison · comparing score against per-task cost
Training–Inference Consistency · train–inference consistency
Unique Run Identifiers · unique ID
Closed-Model Performance with Refusal Substitution · combined performance using closed models
Removing Pattern-Matching-Based Anti-Cheat Checks · removing pattern-matching-based checks
unfiled 231
Gated DeltaNet and Qwen Sparse Attention hybrid architecture · GDN + QSA hybrid architecture
Adaptive Rate Interleave Alignment (ARIA) · ARIA (Adaptive Rate Interleave Alignment)
Blockwise candidate-pool selection for sparse indexing · blockwise candidate selection
disk-based context caching · hard disk caching
Gated DeltaNet and Qwen Sparse Attention hybrid · GDN + QSA
Global attention with a dense feed-forward network in the first Transformer block · global attention with a dense FFN
Hybrid fast-and-slow-thinking model · hybrid fast-and-slow-thinking
IndexShare for multi-token prediction speculative decoding · IndexShare MTP
MXFP4 weight storage with block size 32 · block size of 32
native multimodal training from the start of training · native multimodal training
post-training scaling · scaled post-training
Short convolution (sconv) modules · Short convolution (sconv)
OpenAI-compatible API · OpenAI-compatible
configurable video frame sampling with fps and do_sample_frames · fps=2 and do_sample_frames=True
Omit residual branch-mixing operator · removing the mixing operator
PLE CPU offload for N-gram embeddings · PLE CPU offload
Terminus 2 · Terminus-2 framework
adequate output length (32,768 tokens default, 81,920 for hard benchmarks) · Adequate Output Length
Amazon Elastic Container Service · AWS ECS
Apache 2.0 License · Apache 2.0
archipelago · archipelago codebase
benchmarking against public frontier models · benchmarked against public frontier models
chat templates for model prompting · model chat templates
Configurable routing profiles for model selection · configurable routing profiles
contrastive pre-training of vision encoder · contrastive pre-training
cross-provider failover load balancing · routing to another healthy provider
Cross-provider failover routing for inference requests · routing to another healthy provider
Data quality and diversity enhancement · enhanced data quality and diversity
Failover routing across providers · routing to another healthy provider
Failover routing to healthy upstream providers · routing to another healthy provider
Fine-tuning hyperparameter verification · Fine-tuning performance verification
Fine-tuning via Tinker platform · fine-tune themselves through Tinker
Fixing dependency drift in benchmark tasks · Fixed dependency drift on tasks
Fixing verifier selections in benchmark tasks · fixing verifier selections
four-stage Talker training pipeline · four-stage training pipeline for Talker
Fuse gated residual reads and writes into single kernels · fused into a single kernel
Ignoring known teardown flakes in evaluation · we ignore it
In-task retries for flaky external services · in-task retries
Inference-time scaling plots for evaluation reporting · Inference-time scaling plots
Large collection of reinforcement-learning task environments · more than 7,000 RL task environments
likelihood-based acceptance-rate loss (LK loss) · LK loss for draft model fine-tuning
Limit maximum concurrent sequences to 256 · keep --max-num-seqs 256
masking-based refinement · mask fine-tuning
Max-pool teacher attention distributions to block level · max pooling
Measure per-block maximum activation · per-block maximum activation
multimodal pre-training · multimodal pre-training at scale
ngspice simulation · ngspice simulation loop
Open Model Data Warehouse License Agreement 1.1 · OpenMDW License Agreement, version 1.1
open-weight model release · open-weight systems
Open-weights release under Apache-2.0 license · open-weights model
OpenAI-compatible API endpoint · OpenAI-compatible endpoint
OpenAI-compatible message encoding without Jinja template · encoding folder
Optimization on real agent-harness environments · Optimized on real harness environments
pass@1 metric · pass@1
per-model token-per-second rescaling of timeouts · per-model TPS rescaling
Pinning dependency versions to avoid drift · pinned to setuptools^=58.0.0
Pre-commit hooks with ruff · pre-commit hooks
Pre-training from scratch · pre-trained Inkling from scratch
preview-first release approach · preview-first approach
private benchmark for evaluation · private benchmark
Publication of unedited evaluation trajectories · three unedited trajectories
PyTorch expandable segments allocator (PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True) for loading checkpoints · expandable allocator
Qwen3 reasoning parser · reasoning parser qwen3
recommended sampling parameters · sampling parameters
recommended sampling parameters for deployment · sampling parameters
recommended sampling parameters for thinking and non-thinking modes · Sampling Parameters
Reproducible evaluation recipes in NeMo Gym · evaluation recipes
Request-filter-constrained provider routing · request filters
Run ShellCheck at info severity to catch possible misspellings · Run ShellCheck at info level
shadow replicas for shared indexers in pipeline-parallel training · Shadow indexers
shared-memory caching of preprocessed multimodal inputs · --mm-processor-cache-type shm
Single multi-node Slurm launch · single multi-node srun
staged release · staged approach to release
Text-only multimodal benchmark alignment for comparability · Multimodal benchmark alignment
three-stage SWE teacher training pipeline · three-stage pipeline
tool augmentation during evaluation · tool augmentation
tool-augmented benchmark evaluation (with and without tools) · tool augmentation
Two-stage contextual-parallel communication for compressed attention · two-stage communication approach
WSD optimal learning rate scaling law · WSD optimal-LR scaling law
cross-provider failover routing · routing to another healthy provider
Adequate output length recommendation · output length of 32,768 tokens
AITER Linear backend · VLLM_ROCM_USE_AITER_LINEAR=1
AITER MHA backend · VLLM_ROCM_USE_AITER_MHA=1
AITER RMSNorm backend · VLLM_ROCM_USE_AITER_RMSNORM=1
Balanced routing: price + speed per request · Balanced (price + speed)
Configurable video frame sampling (fps/do_sample_frames) during inference · video frame sampling
Fine-tuning on Tinker platform · availability on Tinker for fine-tuning
Multi-precision weight formats (BF16/FP8/INT4/NVFP4) · weights in BF16, FP8, INT4, and NVFP4
Nitro routing: fastest provider per request · Nitro (fastest)
OpenAI Harmony renderer (openai_harmony) that renders messages into tokens · harmony renderer library
Output length recommendation · output length of 32,768 tokens
Recommended API parameter settings · Recommended Settings
reducing CUDA graph capture size to fit Mamba cache · Reduce --max-cudagraph-capture-size
request routing modes · routing mode
Serving Qwen3.5 with SGLang · SGLang serving framework
Serving Qwen3.5 with vLLM · vLLM serving engine
setting video preprocessor longest_edge to 469,762,048 for hour-scale video understanding · longest_edge parameter
Interaction models (AI that listens, speaks, and interrupts) · interaction models
Tiered CPU and filesystem offloading · Validation used TieringOffloadingSpec
cache-hit cost modeling for agent workloads · model the cache-hit price first
Defense-in-depth for safety · defense-in-depth
Input/output classification with moderation tools · input/output classification
Planning with the model for self-contained tasks · planning with the model
Pre-training methods · New pre-training methods
Sampling with temperature=1.0 and top_p=1.0 · temperature=1.0 and top_p=1.0
AITER fused-MoE backend · VLLM_ROCM_USE_AITER_MOE=1
Read only the two highest-gated residual branches · sparse writes