Model techniques map
Techniquesinference & servingserving parallelism

implementation detail · filed under inference & serving

Token migration for balanced expert placement

Moves tokens from overloaded ranks to underloaded ranks until the target balanced load is reached.

source
1
model
1
labs adopt it
0
strongest
unclear

How sources treat it

One count per evidence span, weakest treatment to strongest.

unclear 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

We repeatedly pick an underloaded rank and an overloaded rank, and migrate tokens from the overloaded rank to fill the underloaded rank exactly up to the balanced value S×K

unclearmodel architecturein Kimi K3Moonshot AI

Filed alongside

Other methods under inference & serving :: serving parallelism.