implementation detail · filed under model architecture
10 Routed + 1 Shared Experts Activated per Token
A reported MoE activation configuration with 10 routed experts and 1 shared expert activated per token.
Also called 10 routed + 1 shared activated per token, 10 Routed + 1 Shared.
- sources
- 2
- models
- 2
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
Number of Activated Experts: 10 Routed + 1 Shared
coremodel architecturein Qwen3.5-397B-A17BQwen
MoE experts: 512 experts, 10 routed + 1 shared activated per token
coremodel architecturein Qwen3.8-Flash-NextQwen
Filed alongside
Other methods under model architecture :: channel mixer :: mixture of experts :: shared experts.