taxonomy node · level 4
expert routing
15 methods filed at this node or below it, from the sources of 14 models.
model architecture :: channel mixer :: mixture of experts :: expert routing
Matching aids for the classifier: top-k routing; gating; hash routing; anticipatory routing; latent-space routing; routed experts.
In this branch 15
Everything filed at this node or below it, with one collapsible heading per child node.
filed here 15
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| GLM-5.3-Flash core | Token-level expert routing core— |
| Hy4-preview core | Top-8 expert routing core— |
| DeepSeek-V4-Flash used | Hash routing usedAnticipatory Routing optionalLoss-spike-triggered Anticipatory Routing optional— |
| MiMo-V2.5 used | Top-8 expert routing used— |
| Hy3 core | Routed experts coreTop-8 expert routing core— |
| GLM-5.2 used | DP-aware routing used— |
| MiniMax-M3 used | Top-4 expert routing used— |
| DeepSeek-V3.2 core | Keep Routing core— |
| DeepSeek-V4-Pro used | Anticipatory Routing usedHash routing usedLoss-spike-triggered Anticipatory Routing optional— |
| Inkling core | Token-level expert routing coreTop-6 expert routing core— |
| Kimi K3 default | Fixed Top-k routing with frozen bias defaultLatent-space routing used— |
| Laguna-S-2.1 core | Token-choice routing with softplus gating coreToken-choice routing used— |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B core | Token-level expert routing core— |
| gpt-oss-120b core | Top-k expert routing with softmax over selected experts core— |