Model techniques map
Techniquesoptimizationlearning-rate schedule

specific method · filed under optimization

Warmup-Stable-Decay

A learning-rate schedule identified as WSD, distinct in the evidence from cosine decay.

Also called Warmup Stable Decay learning rate schedule, Warmup Stable Decay (WSD).

sources
4
models
3
labs adopt it
2
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 3not used 2

Documented in

Further reading

Picked by hand, not extracted: where to read more, not evidence for anything on this page.

Evidence

5 spans quoted from the sources, strongest treatment first.

Warmup-Stable-Decay (WSD) learning rate schedule over a total horizon of 20 trillion tokens

usedoptimizationin Nemotron 3 UltraNVIDIA

we adopted a Warmup-Stable-Decay (WSD) learning rate schedule instead of cosine.

usedoptimizationin Laguna XS.2Poolside

we use a Warmup-Stable-Decay (WSD) learning rate schedule

usedoptimizationin Nemotron 3 UltraNVIDIA

cosine decay consistently achieves a lower final loss than WSD.

not usedoptimizationin Kimi K3Moonshot AI

Our scaling-law study consistently favors cosine decay over Warmup Stable Decay (WSD), leading us to adopt cosine decay as the default learning rate schedule.

not usedunclearin Kimi K3Moonshot AI

Filed alongside

Other methods under optimization :: learning-rate schedule.