specific method · filed under optimization
Category-specific assignment of Muon and AdamW
A training recipe that assigns Muon and AdamW to different parameter categories.
Also called Tailored Training Recipe, division of labour between Muon and AdamW.
- sources
- 3
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 3core 1
Documented in
Evidence
4 spans quoted from the sources, strongest treatment first.
The Muon and AdamW optimizers are applied to specific weight categories to maximize efficiency.
coreunclearin Qwen3.8-Flash-NextQwen
Muon and AdamW applied to specific weight categories
usedoptimizationin Qwen3.8-Flash-NextQwen
The Muon and AdamW optimizers are applied to specific weight categories
usedunclearin Qwen3.8-Flash-NextQwen
the division of labour between Muon and AdamW
usedoptimizationin Qwen3.8-Flash-NextQwen
Filed alongside
Other methods under optimization :: optimizer.
MuonPer-Head MuonAdamWSplit fused gradients before orthogonalizationNesterov momentumSinkhorn-balanced updateAdam without weight decay for the N-gram embedding tableAdamW for attention and GDN output gatesAdamW for gated-residual low-rank projectionsAdamW for input embeddings and output headAdamW for the MoE routerAsynchronous Micro-Group pipelineCanzonaCUDA graph capture of the optimizer stepEight-step Newton–Schulz iterationHybrid Muon and AdamW parameter-group optimizer assignmentHybrid Newton–Schulz iterationsHybrid optimization with Muon and AdamMoonlight-style learning-rate scalingMuon orthogonalization accuracy refinementMuon restricted to two-dimensional linear-map weightsMuon SplitMuownPeer-to-peer shard retrieval for Muon orthogonalization