implementation detail · filed under model architecture
Top-8 expert routing
Selects eight routed experts for each token.
Also called 192 experts, top-8 activated, top-8 routed experts, Top-8 expert selection, top-8, Top-8 routing, 256 routed experts (top-8).
- sources
- 5
- models
- 3
- labs adopt it
- 2
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2core 3
Documented in
Evidence
5 spans quoted from the sources, strongest treatment first.
Number of Experts | 192 experts, top-8 activated
coremodel architecturein Hy3Tencent Hunyuan
every token activates the top-8 routed experts along with the shared expert
coremodel architecturein Hy4-previewTencent Hunyuan
top-8
coremodel architecturein Hy4-previewTencent Hunyuan
it uses 256 routed experts (top-8)
usedmodel architecturein MiMo-V2.5Xiaomi
Hy3’s architecture contains a sparse MoE with 192 experts and top-8 routing.
usedmodel architecturein Hy3Tencent
Filed alongside
Other methods under model architecture :: channel mixer :: mixture of experts :: expert routing.
Token-level expert routingAnticipatory RoutingTop-4 expert routingDP-aware routingFixed Top-k routing with frozen biasHash routingKeep RoutingLatent-space routingLoss-spike-triggered Anticipatory RoutingRouted expertsToken-choice routingToken-choice routing with softplus gatingTop-6 expert routingTop-k expert routing with softmax over selected experts