Model techniques map
Techniquesoptimizationoptimizer

specific method · filed under optimization

AdamW for attention and GDN output gates

Using AdamW for output gates where ablations found it comparable to or slightly better than Muon.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

For the output gates (the attention output gate ... and the GDN projection), our ablations found AdamW on par with or slightly better than Muon.

usedoptimizationin Qwen3.8-NextQwen

Filed alongside

Other methods under optimization :: optimizer.