implementation detail · filed under model architecture
Expert Parallelism Load Balancing
Balances expert-parallel execution by replicating hot experts across expert-parallel ranks.
Also called EPLB.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
replicating hot experts across EP ranks (EPLB)
usedinference servingin Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under model architecture :: channel mixer :: mixture of experts :: expert load balancing.
Auxiliary-loss-free load balancingQuantile BalancingHistogram-based quantile estimationAuxiliary-loss load balancingExact coordinate minimizationExpert bias update factorExponential moving average of estimated quantilesformHCLoad balancingMaxVioModality-specific auxiliary-loss-free load balancingPersistent load balancingPooled-global-batch quantile estimationRound-robin load balancing