implementation detail · filed under optimization
Nesterov momentum
Applying Nesterov momentum in the Muon update or orthogonalization process.
Also called Nesterov-accelerated momentum, Nesterov trick.
- sources
- 2
- models
- 3
- labs adopt it
- 2
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
we also apply weight decay to Muon parameters, use the Nesterov trick
usedoptimizationin DeepSeek-V4DeepSeek
The NS iteration is applied to a Nesterov-accelerated momentum
usedoptimizationin Qwen3.8-NextQwen
Filed alongside
Other methods under optimization :: optimizer.
MuonPer-Head MuonAdamWCategory-specific assignment of Muon and AdamWSplit fused gradients before orthogonalizationSinkhorn-balanced updateAdam without weight decay for the N-gram embedding tableAdamW for attention and GDN output gatesAdamW for gated-residual low-rank projectionsAdamW for input embeddings and output headAdamW for the MoE routerAsynchronous Micro-Group pipelineCanzonaCUDA graph capture of the optimizer stepEight-step Newton–Schulz iterationHybrid Muon and AdamW parameter-group optimizer assignmentHybrid Newton–Schulz iterationsHybrid optimization with Muon and AdamMoonlight-style learning-rate scalingMuon orthogonalization accuracy refinementMuon restricted to two-dimensional linear-map weightsMuon SplitMuownPeer-to-peer shard retrieval for Muon orthogonalization