Model techniques map
Taxonomyoptimizationoptimizer

taxonomy node · level 2

optimizer

29 methods filed at this node or below it, from the sources of 11 models.

optimization :: optimizer

Matching aids for the classifier: AdamW; Muon; per-head Muon; Muon Split; Nesterov momentum; Newton-Schulz iterations.

In this branch 29

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 29

Muon core · 11 sources · 13 quotes
AdamW for the MoE router core · 1 source · 1 quote
Per-Head Muon default · 4 sources · 5 quotes
Eight-step Newton–Schulz iteration default · 1 source · 1 quote
Sinkhorn-balanced update default · 1 source · 2 quotes
AdamW used · 3 sources · 4 quotes
Nesterov momentum used · 2 sources · 2 quotes
AdamW for attention and GDN output gates used · 1 source · 1 quote
AdamW for input embeddings and output head used · 1 source · 1 quote
Asynchronous Micro-Group pipeline used · 1 source · 1 quote
Canzona used · 1 source · 1 quote
CUDA graph capture of the optimizer step used · 1 source · 1 quote
Hybrid Newton–Schulz iterations used · 1 source · 1 quote
Hybrid optimization with Muon and Adam used · 1 source · 1 quote
Moonlight-style learning-rate scaling used · 1 source · 1 quote
Muon orthogonalization accuracy refinement used · 1 source · 1 quote
Muon Split used · 1 source · 1 quote
Muown used · 1 source · 1 quote
Polar Express coefficient schedule used · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modeltechniques
DeepSeek-V4.1-Flash defaultPer-Head Muon defaultSinkhorn-balanced update defaultAdamW usedMuon used—
MiMo-V2.6-Flash usedMuown used—
DeepSeek-V4-Flash coreMuon coreAdamW usedHybrid Newton–Schulz iterations usedNesterov momentum used—
MiMo-V2.5 usedAdamW used—
GLM-5.2 usedMuon usedMuon Split used—
DeepSeek-V4-Pro coreMuon coreAdamW usedHybrid Newton–Schulz iterations usedNesterov momentum used—
Inkling usedHybrid optimization with Muon and Adam used—
Kimi K3 usedMuon usedPeer-to-peer shard retrieval for Muon orthogonalization usedPer-Head Muon used—
Laguna-S-2.1 coreMuon coreMoonlight-style learning-rate scaling used—
MiMo-V2.6-Pro usedMuown used—
Qwen3.8-Flash-Next coreAdamW for the MoE router coreCategory-specific assignment of Muon and AdamW coreMuon coreMuon restricted to two-dimensional linear-map weights coreSplit fused gradients before orthogonalization coreEight-step Newton–Schulz iteration defaultHybrid Muon and AdamW parameter-group optimizer assignment defaultAdam without weight decay for the N-gram embedding table usedAdamW for attention and GDN output gates usedAdamW for gated-residual low-rank projections usedAdamW for input embeddings and output head usedAsynchronous Micro-Group pipeline usedCanzona usedCUDA graph capture of the optimizer step usedMuon orthogonalization accuracy refinement usedNesterov momentum usedPer-head splitting of attention and GDN input projections usedPolar Express coefficient schedule usedSelected number of Newton–Schulz iterations usedSplit SwiGLU fc1 into gate and up sub-matrices used—