specific method · filed under optimization
Weight decay coupled to learning-rate squared
Sets weight-decay strength proportional to the square of the learning rate to keep model-weight size stable across training horizons.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
We coupled the weight decay strength to the square of the learning rate, which we found kept the overall size of the model weights stable across training horizons
usedoptimizationin InklingThinking Machines Lab
Filed alongside
Other methods under optimization :: training stability.
Weight paddingActivation or logit clippingAnti-hallucination trainingBehavioral regularizationCross-replica model-weight hash consistency checksElevated constant-learning-rate training stability stress testExponential moving average of checkpointsGradient clippingLearning-rate elevation for training stability stress testingLoss masking for excessively stale tokensOff-policy sample filteringPer-token regularization for off-policy RLSafeguards against training drift and reward hackingSoft droppingTraining stability stress testingWeight clipping