Model techniques map
Taxonomyoptimization

taxonomy area · pipeline stage 04

optimization

155 methods filed at this node or below it, from the sources of 22 models.

optimization

Matching aids for the classifier: how the training run is made to converge and stay stable.

In this branch 155

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 4

Updated hyperparameter scaling law core · 3 sources · 4 quotes
Momentum update with Sinkhorn balancing core · 1 source · 1 quote
Fused loss computation used · 1 source · 1 quote
Phase-by-phase training cost accounting used · 1 source · 1 quote

optimizer 29

Muon core · 11 sources · 13 quotes
AdamW for the MoE router core · 1 source · 1 quote
Per-Head Muon default · 4 sources · 5 quotes
Eight-step Newton–Schulz iteration default · 1 source · 1 quote
Sinkhorn-balanced update default · 1 source · 2 quotes
AdamW used · 3 sources · 4 quotes
Nesterov momentum used · 2 sources · 2 quotes
AdamW for attention and GDN output gates used · 1 source · 1 quote
AdamW for input embeddings and output head used · 1 source · 1 quote
Asynchronous Micro-Group pipeline used · 1 source · 1 quote
Canzona used · 1 source · 1 quote
CUDA graph capture of the optimizer step used · 1 source · 1 quote
Hybrid Newton–Schulz iterations used · 1 source · 1 quote
Hybrid optimization with Muon and Adam used · 1 source · 1 quote
Moonlight-style learning-rate scaling used · 1 source · 1 quote
Muon orthogonalization accuracy refinement used · 1 source · 1 quote
Muon Split used · 1 source · 1 quote
Muown used · 1 source · 1 quote
Polar Express coefficient schedule used · 1 source · 1 quote

learning-rate schedule 13

Cosine decay default · 4 sources · 5 quotes
Warmup-Stable-Decay used · 4 sources · 5 quotes
Batch-size warmup used · 2 sources · 4 quotes
Training without batch-size warmup used · 2 sources · 2 quotes
Engram learning-rate scaling used · 1 source · 1 quote
Fixed elevated learning rate used · 1 source · 1 quote
Learning-rate annealing used · 1 source · 1 quote
Linear warmup followed by cosine decay used · 1 source · 1 quote
Scaling-law fit used · 1 source · 1 quote
Scheduled batch-size growth used · 1 source · 1 quote
WSD-specific scaling law used · 1 source · 1 quote

training precision 19

NVFP4 core · 10 sources · 13 quotes
NVFP4 pre-training core · 5 sources · 5 quotes
BF16 mixed-precision training default · 1 source · 1 quote
BF16 used · 5 sources · 5 quotes
FP8 mixed-precision training used · 4 sources · 4 quotes
FP4+FP8 mixed precision used · 2 sources · 2 quotes
FP8-precision reinforcement learning used · 2 sources · 2 quotes
MXFP8 used · 2 sources · 2 quotes
E2M1 used · 1 source · 1 quote
FP32 attention-output retention used · 1 source · 1 quote
FP32 gradient reduction used · 1 source · 1 quote
FP8 storage for the residual state used · 1 source · 1 quote
High-precision final network layers used · 1 source · 1 quote
Mixed-FP8 quantization used · 1 source · 1 quote
Mixed-precision training used · 1 source · 1 quote
NVFP4 fine-grained micro-block scaling used · 1 source · 1 quote
Two-dimensional block quantization used · 1 source · 1 quote
BF16 gradient reduction not used · 1 source · 1 quote

quantization-aware training 15

Quantization-Aware Training core · 8 sources · 8 quotes
Quantize-Dequantize Training core · 1 source · 1 quote
FP4 Quantization default · 4 sources · 4 quotes
FP4 Quantization-Aware Training used · 3 sources · 3 quotes
MXFP4 Weights with MXFP8 Activations used · 2 sources · 3 quotes
Stochastic Rounding used · 2 sources · 3 quotes
INT4 Quantization-Aware Training used · 1 source · 1 quote
MXFP4 Quantization-Aware Post-Training used · 1 source · 1 quote
NVFP4 Training used · 1 source · 1 quote
Per-Block Scalar Scaling used · 1 source · 1 quote
Q4_0 Quantization Format used · 1 source · 1 quote
Random Hadamard Transforms used · 1 source · 1 quote
Stochastic Rounding for Mamba Cache used · 1 source · 1 quote
Stochastic Rounding of Gradients used · 1 source · 1 quote
Straight-Through Estimator used · 1 source · 1 quote

training stability 17

Weight padding used · 2 sources · 2 quotes
Anti-hallucination training used · 1 source · 1 quote
Behavioral regularization used · 1 source · 1 quote
Exponential moving average of checkpoints used · 1 source · 1 quote
Gradient clipping used · 1 source · 1 quote
Loss masking for excessively stale tokens used · 1 source · 1 quote
Off-policy sample filtering used · 1 source · 1 quote
Per-token regularization for off-policy RL used · 1 source · 1 quote
Soft dropping used · 1 source · 1 quote
Training stability stress testing used · 1 source · 1 quote
Weight clipping used · 1 source · 1 quote
Activation or logit clipping not used · 1 source · 1 quote

training parallelism 29

Communication-Computation Overlap core · 1 source · 2 quotes
MoonEP core · 1 source · 4 quotes
Redundant-Expert Capacity Reservation core · 1 source · 1 quote
Sequence Parallelism for Activations core · 1 source · 1 quote
Expert Parallelism used · 10 sources · 13 quotes
Tensor Parallelism used · 5 sources · 6 quotes
Pipeline Parallelism used · 2 sources · 2 quotes
Cache-Based Pipeline Communication used · 1 source · 1 quote
Context Parallelism used · 1 source · 1 quote
Fully Balanced Expert-Parallel Training used · 1 source · 1 quote
Hybrid ZeRO Bucket Assignment for Muon used · 1 source · 1 quote
KDA Context Parallelism used · 1 source · 1 quote
Load-Balanced Image Sharding used · 1 source · 1 quote
Pipeline Payload Extensions used · 1 source · 1 quote
SConv-Aware Tensor-Parallel Sharding used · 1 source · 1 quote
SM-Level Context Parallelism used · 1 source · 1 quote
ZeRO-3 Optimizer State Sharding used · 1 source · 1 quote
Data-Weighted Data Parallelism mentioned · 1 source · 1 quote

training runtime 29

Composable activation storage policies core · 1 source · 1 quote
In-flight recovery system core · 1 source · 1 quote
Offline checkpoint merging used · 3 sources · 3 quotes
Asynchronous checkpointing used · 2 sources · 2 quotes
Autotune configuration generation used · 1 source · 1 quote
Batch-level embedding prefetch used · 1 source · 1 quote
CPU-resident optimizer states used · 1 source · 1 quote
Cross-rank remote activation offloading used · 1 source · 1 quote
Element-wise activation recomputation used · 1 source · 1 quote
Heartbeat-driven fault tolerance used · 1 source · 1 quote
In-flight weight updates used · 1 source · 1 quote
Job-level eviction and reclaim used · 1 source · 1 quote
Persistent checkpoint worker processes used · 1 source · 1 quote
Pre-admission hardware stress testing used · 1 source · 1 quote
Same-node sticky pod respawn used · 1 source · 1 quote
Slice-Granularity Elasticity used · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modelfiled hereoptimizerlearning-rate scheduletraining precisionquantization-aware trainingtraining stabilitytraining parallelismtraining runtime
DeepSeek-V4.1-Flash coreMomentum update with Sinkhorn balancing core—Per-Head Muon defaultSinkhorn-balanced update defaultAdamW usedMuon used—Engram learning-rate scaling usedLinear warmup followed by cosine decay used———Loss masking for excessively stale tokens used—Communication-Computation Overlap coreLoad-Balanced Image Sharding usedPipeline Payload Extensions used—Batch-level embedding prefetch usedOffline checkpoint merging used—
NVIDIA-Nemotron-3-Ultra-550B-A55B core——Cosine decay usedLearning-rate annealing usedWarmup-Stable-Decay used—NVFP4 coreNVFP4 pre-training coreBF16 usedE2M1 usedFP32 gradient reduction usedMixed-FP8 quantization usedMXFP8 usedTwo-dimensional block quantization usedBF16 gradient reduction not used—Stochastic Rounding usedStochastic Rounding for Mamba Cache used—Weight padding used—Context Parallelism usedExpert Parallelism usedPipeline Parallelism usedTensor Parallelism usedData-Weighted Data Parallelism mentioned—Asymmetric local-read and single-writer cache paths usedAsynchronous checkpointing usedCaching the distributed checkpoint save plan usedIn-flight weight updates usedOffline checkpoint merging usedOverlapping NCCL transfers with device-to-host copies usedPersistent checkpoint worker processes usedPersistent shared-storage cache for compiled artifacts usedSeeding node-local storage from a warm shared cache used—
MiMo-V2.6-Flash coreFused loss computation used—Muown used———Quantize-Dequantize Training coreFP4 Quantization-Aware Training used—Behavioral regularization usedSafeguards against training drift and reward hacking used—Window-Bounded KV Exchange under Context Parallelism used—CPU-resident optimizer states used—
DeepSeek-V4-Flash core—Muon coreAdamW usedHybrid Newton–Schulz iterations usedNesterov momentum used—Scheduled batch-size growth used—FP4+FP8 mixed precision used—FP4 Quantization usedFP4 Quantization-Aware Training usedQuantization-Aware Training usedStraight-Through Estimator used——All-to-all Gradient Exchange with Local FP32 Summation usedHybrid ZeRO Bucket Assignment for Muon usedKnapsack-Based Balanced Assignment of Dense Parameter Matrices usedModified DualPipe 1F1B Pipeline Overlap for mHC used—Tensor-level automatic-differentiation activation checkpointing coreMinimal recomputation-graph extraction by backward traversal usedSelective tensor recomputation for activation-memory reduction used—
MiMo-V2.5 used—AdamW used—Cosine decay used—FP8 mixed-precision training used——Gradient clipping used———
GLM-5.3 used——————Expert Parallelism used——
Hy3 used—————Anti-hallucination training used———
GLM-5.2 used—Muon usedMuon Split used—Batch-size warmup usedCosine decay used—Mixed-precision training used—INT4 Quantization-Aware Training used—Off-policy sample filtering used—Expert Parallelism used—Heartbeat-driven fault tolerance used—
MiniMax-M3 used———————Autotune configuration generation used—
DeepSeek-V4-Pro core—Muon coreAdamW usedHybrid Newton–Schulz iterations usedNesterov momentum used—Scheduled batch-size growth used—FP4+FP8 mixed precision used—FP4 Quantization usedFP4 Quantization-Aware Training usedQuantization-Aware Training usedStraight-Through Estimator used——All-to-all Gradient Exchange with Local FP32 Summation usedHybrid ZeRO Bucket Assignment for Muon usedKnapsack-Based Balanced Assignment of Dense Parameter Matrices usedModified DualPipe 1F1B Pipeline Overlap for mHC used—Tensor-level automatic-differentiation activation checkpointing coreMinimal recomputation-graph extraction by backward traversal usedSelective tensor recomputation for activation-memory reduction used—
Gemma 4 31B core————Quantization-Aware Training corePer-Block Scalar Scaling usedQ4_0 Quantization Format used——Data Replica Reduction over the Data Center Network usedZeRO-3 Optimizer State Sharding used—Slice-Granularity Elasticity used—
Inkling used—Hybrid optimization with Muon and Adam used——NVFP4 optional——Weight decay coupled to learning-rate squared used—SConv-Aware Tensor-Parallel Sharding used——
Kimi K3 coreUpdated hyperparameter scaling law used—Muon usedPeer-to-peer shard retrieval for Muon orthogonalization usedPer-Head Muon used—Cosine decay defaultWarmup-Stable-Decay not used—Block-wise FP8 activation quantization with offload usedFP32 attention-output retention used—MXFP4 Quantization-Aware Post-Training usedMXFP4 Weights with MXFP8 Activations usedQuantization-Aware Training used—Per-token regularization for off-policy RL usedSoft dropping usedWeight clipping used—MoonEP coreRedundant-Expert Capacity Reservation coreSequence Parallelism for Activations coreCache-Based Pipeline Communication usedDynamic Context Parallelism for Large Multimodal Samples usedExpert Parallelism usedFully Balanced Expert-Parallel Training usedGPU Planning Kernel for Redundant Expert Migration usedKDA Context Parallelism usedPipeline ZeRO-2 Gradient Sharding with CPU Offloading usedpipeline-bubble scheduling of ViT computation usedSM-Level Context Parallelism used—Composable activation storage policies coreUnified pluggable activation storage manager coreCross-rank remote activation offloading usedDouble-buffered chunked streaming of reference-model weights usedElement-wise activation recomputation usedPolicy-model gradient-buffer reuse for reference-model weights usedPost-iteration NVMe offloading of training states used—
Laguna-S-2.1 core—Muon coreMoonlight-style learning-rate scaling used—Warmup-Stable-Decay usedWSD-specific scaling law used—BF16 mixed-precision training defaultFP8-precision reinforcement learning used——Cross-replica model-weight hash consistency checks usedExponential moving average of checkpoints used——In-flight recovery system coreGPU-to-GPU weight synchronization over GPUDirect RDMA usedJob-level eviction and reclaim usedPre-admission hardware stress testing usedSame-node sticky pod respawn used—
MiMo-V2.5-Pro used———FP8 mixed-precision training used—————
MiMo-V2.6-Pro coreFused loss computation used—Muown used———Quantize-Dequantize Training coreFP4 Quantization-Aware Training used—Behavioral regularization usedSafeguards against training drift and reward hacking used—Window-Bounded KV Exchange under Context Parallelism used—CPU-resident optimizer states used—
NVIDIA-Nemotron-3.5-Lightning-30B-A3B core———NVFP4 coreBF16 usedHigh-precision final network layers usedMXFP8 usedNVFP4 fine-grained micro-block scaling usedNVFP4 pre-training used—NVFP4 Training usedRandom Hadamard Transforms usedStochastic Rounding of Gradients used————
Qwen3.5-397B-A17B used——————Expert Parallelism used——
Qwen3.6-35B-A3B used——————Tensor Parallelism used——
Qwen3.8-Flash-Next coreUpdated hyperparameter scaling law corePhase-by-phase training cost accounting used—AdamW for the MoE router coreCategory-specific assignment of Muon and AdamW coreMuon coreMuon restricted to two-dimensional linear-map weights coreSplit fused gradients before orthogonalization coreEight-step Newton–Schulz iteration defaultHybrid Muon and AdamW parameter-group optimizer assignment defaultAdam without weight decay for the N-gram embedding table usedAdamW for attention and GDN output gates usedAdamW for gated-residual low-rank projections usedAdamW for input embeddings and output head usedAsynchronous Micro-Group pipeline usedCanzona usedCUDA graph capture of the optimizer step usedMuon orthogonalization accuracy refinement usedNesterov momentum usedPer-head splitting of attention and GDN input projections usedPolar Express coefficient schedule usedSelected number of Newton–Schulz iterations usedSplit SwiGLU fc1 into gate and up sub-matrices used—Fixed elevated learning rate usedScaling-law fit usedTraining without batch-size warmup usedBatch-size warmup with adjusted peak learning rate evaluatedBatch-size warmup with constant-batch optimum peak learning rate evaluatedBatch-size warmup not used—FP8 storage for the residual state used——Elevated constant-learning-rate training stability stress test usedLearning-rate elevation for training stability stress testing usedTraining stability stress testing usedActivation or logit clipping not used—NS-FLOP-Balanced Static Parameter Partitioning usedTensor and Expert Parallelism with Degree Eight used——
Step-3.7-Flash used———BF16 usedNVFP4 optional———Expert Parallelism usedTensor Parallelism used——
gpt-oss-120b default————FP4 Quantization default————