specific method · filed under optimization
Training without batch-size warmup
A strategy that skips batch-size warmup and starts training at the target batch size.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
the elimination of traditional batch-size warmups in favour of starting directly at the target batch size
usedoptimizationin Qwen3.8-Flash-NextQwen
we eliminate traditional batch-size warmups and start directly at the target batch size
usedunclearin Qwen3.8-Flash-NextQwen
Filed alongside
Other methods under optimization :: learning-rate schedule.
Cosine decayWarmup-Stable-DecayBatch-size warmupBatch-size warmup with adjusted peak learning rateBatch-size warmup with constant-batch optimum peak learning rateEngram learning-rate scalingFixed elevated learning rateLearning-rate annealingLinear warmup followed by cosine decayScaling-law fitScheduled batch-size growthWSD-specific scaling law