taxonomy node · level 2
learning-rate schedule
13 methods filed at this node or below it, from the sources of 9 models.
optimization :: learning-rate schedule
Matching aids for the classifier: warmup-stable-decay; WSD; cosine decay; learning rate annealing; batch size warmup.
In this branch 13
Everything filed at this node or below it, with one collapsible heading per child node.
filed here 13
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| DeepSeek-V4.1-Flash used | Engram learning-rate scaling usedLinear warmup followed by cosine decay used— |
| NVIDIA-Nemotron-3-Ultra-550B-A55B used | Cosine decay usedLearning-rate annealing usedWarmup-Stable-Decay used— |
| DeepSeek-V4-Flash used | Scheduled batch-size growth used— |
| MiMo-V2.5 used | Cosine decay used— |
| GLM-5.2 used | Batch-size warmup usedCosine decay used— |
| DeepSeek-V4-Pro used | Scheduled batch-size growth used— |
| Kimi K3 default | Cosine decay defaultWarmup-Stable-Decay not used— |
| Laguna-S-2.1 used | Warmup-Stable-Decay usedWSD-specific scaling law used— |
| Qwen3.8-Flash-Next used | Fixed elevated learning rate usedScaling-law fit usedTraining without batch-size warmup usedBatch-size warmup with adjusted peak learning rate evaluatedBatch-size warmup with constant-batch optimum peak learning rate evaluatedBatch-size warmup not used— |