general family · filed under optimization
Updated hyperparameter scaling law
A scaling law is updated or refitted to retune training hyperparameters such as learning rate and batch size for a new architecture or optimizer.
Also called refit the scaling law, dedicated scaling-law studies.
- sources
- 3
- models
- 2
- labs adopt it
- 2
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Evidence
4 spans quoted from the sources, strongest treatment first.
Therefore, we develop an updated hyperparameter scaling law.
we conduct dedicated scaling-law studies to retune key hyperparameters, including the batch size, learning rate, tokens-per-parameter ratio (TPP) and the model shape.
The new architecture and optimizer also shift the optimal hyperparameters, so we refit the scaling law used for the Qwen3.5 series.
with the scaling law refitted for the new architecture
Filed alongside
Other methods under optimization.