Model techniques map
Techniquesoptimization

general family · filed under optimization

Updated hyperparameter scaling law

A scaling law is updated or refitted to retune training hyperparameters such as learning rate and batch size for a new architecture or optimizer.

Also called refit the scaling law, dedicated scaling-law studies.

sources
3
models
2
labs adopt it
2
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 3core 1

Documented in

Evidence

4 spans quoted from the sources, strongest treatment first.

Therefore, we develop an updated hyperparameter scaling law.

coreoptimizationin Qwen3.8-NextQwen

we conduct dedicated scaling-law studies to retune key hyperparameters, including the batch size, learning rate, tokens-per-parameter ratio (TPP) and the model shape.

usedoptimizationin Kimi K3Moonshot AI

The new architecture and optimizer also shift the optimal hyperparameters, so we refit the scaling law used for the Qwen3.5 series.

usedunclearin Qwen3.8-Flash-NextQwen

with the scaling law refitted for the new architecture

usedoptimizationin Qwen3.8-Flash-NextQwen

Filed alongside

Other methods under optimization.