implementation detail · filed under optimization
Engram learning-rate scaling
Scaling the learning rate used for Engram by a stated factor.
Also called learning rate of Engram is scaled by 5×.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Following (Cheng et al., 2026b), the learning rate of Engram is scaled by 5×.
usedunclearin DeepSeek-V4.1-FlashDeepSeek
Filed alongside
Other methods under optimization :: learning-rate schedule.
Cosine decayWarmup-Stable-DecayBatch-size warmupTraining without batch-size warmupBatch-size warmup with adjusted peak learning rateBatch-size warmup with constant-batch optimum peak learning rateFixed elevated learning rateLearning-rate annealingLinear warmup followed by cosine decayScaling-law fitScheduled batch-size growthWSD-specific scaling law