implementation detail Β· not yet filed
Parabola fit in log10 learning rate space to find optimal learning rate
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
For each (π, π· β {60, 120, 240, 480, 960} B) we fit a parabola in log10(lr) space and take the vertex as lrβ(π, π·)
usedoptimizationin Laguna XS.2Poolside