Model techniques map
Techniquesmodel architecturechannel mixermixture of expertsshared experts

implementation detail · filed under model architecture

10 Routed + 1 Shared Experts Activated per Token

A reported MoE activation configuration with 10 routed experts and 1 shared expert activated per token.

Also called 10 routed + 1 shared activated per token, 10 Routed + 1 Shared.

sources
2
models
2
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 2

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

Number of Activated Experts: 10 Routed + 1 Shared

coremodel architecturein Qwen3.5-397B-A17BQwen

MoE experts: 512 experts, 10 routed + 1 shared activated per token

coremodel architecturein Qwen3.8-Flash-NextQwen

Filed alongside

Other methods under model architecture :: channel mixer :: mixture of experts :: shared experts.