implementation detail · filed under optimization
Batch-size warmup with constant-batch optimum peak learning rate
A batch-size warmup variant that keeps the peak learning rate at the constant-batch optimum.
- source
- 1
- model
- 1
- labs adopt it
- 0
- strongest
- evaluated
How sources treat it
One count per evidence span, weakest treatment to strongest.
evaluated 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
the first maintains the peak learning rate of the constant-batch optimum, making the ramp its only difference
evaluatedoptimizationin Qwen3.8-NextQwen
Filed alongside
Other methods under optimization :: learning-rate schedule.
Cosine decayWarmup-Stable-DecayBatch-size warmupTraining without batch-size warmupBatch-size warmup with adjusted peak learning rateEngram learning-rate scalingFixed elevated learning rateLearning-rate annealingLinear warmup followed by cosine decayScaling-law fitScheduled batch-size growthWSD-specific scaling law