implementation detail · filed under model architecture
Top-6 expert routing
Selects six experts from the available routed experts for each token.
Also called each token is routed to 6 of 256 experts.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
each token is routed to 6 of 256 experts
coremodel architecturein InklingThinking Machines Lab
Filed alongside
Other methods under model architecture :: channel mixer :: mixture of experts :: expert routing.
Top-8 expert routingToken-level expert routingAnticipatory RoutingTop-4 expert routingDP-aware routingFixed Top-k routing with frozen biasHash routingKeep RoutingLatent-space routingLoss-spike-triggered Anticipatory RoutingRouted expertsToken-choice routingToken-choice routing with softplus gatingTop-k expert routing with softmax over selected experts