implementation detail · filed under model architecture
Exponential moving average of estimated quantiles
A load-balancing refinement that averages estimated quantiles across steps to reduce batch-to-batch sampling noise.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
maintaining an exponential moving average of the estimated quantiles across steps reduces batch-to-batch sampling noise and can improve load balance still further.
usedmodel architecturein Kimi K3Moonshot AI
Filed alongside
Other methods under model architecture :: channel mixer :: mixture of experts :: expert load balancing.
Auxiliary-loss-free load balancingQuantile BalancingHistogram-based quantile estimationAuxiliary-loss load balancingExact coordinate minimizationExpert bias update factorExpert Parallelism Load BalancingformHCLoad balancingMaxVioModality-specific auxiliary-loss-free load balancingPersistent load balancingPooled-global-batch quantile estimationRound-robin load balancing