ambiguous · filed under model architecture
Load balancing
A general strategy for assigning prompts to GPUs based on load; the evidence does not specify a particular mechanism.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
we reference historical pass rates and, if necessary, assign new prompts to GPUs with load balancing
usedoptimizationin MiMo-V2-FlashXiaomi
Filed alongside
Other methods under model architecture :: channel mixer :: mixture of experts :: expert load balancing.
Auxiliary-loss-free load balancingQuantile BalancingHistogram-based quantile estimationAuxiliary-loss load balancingExact coordinate minimizationExpert bias update factorExpert Parallelism Load BalancingExponential moving average of estimated quantilesformHCMaxVioModality-specific auxiliary-loss-free load balancingPersistent load balancingPooled-global-batch quantile estimationRound-robin load balancing