taxonomy node · level 4
shared experts
7 methods filed at this node or below it, from the sources of 14 models.
model architecture :: channel mixer :: mixture of experts :: shared experts
Matching aids for the classifier: shared expert; no shared experts.
In this branch 7
Everything filed at this node or below it, with one collapsible heading per child node.
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| DeepSeek-V4.1-Flash used | Shared Experts used— |
| Hy4-preview core | Shared Experts coreTop-8 Routed-Expert Selection with Shared-Expert Activation core— |
| NVIDIA-Nemotron-3-Ultra-550B-A55B used | Shared Experts used— |
| MiMo-V2.6-Flash core | No Shared Experts core— |
| DeepSeek-V4-Flash core | Shared Experts core— |
| MiMo-V2.5 not used | No Shared Experts not used— |
| Hy3 core | Shared Experts core— |
| MiniMax-M3 used | Shared Experts used— |
| DeepSeek-V4-Pro core | Shared Experts core— |
| Inkling core | MoE with Expert Sinks coreShared Experts Active on Every Token core— |
| Laguna-S-2.1 used | Shared Experts used— |
| MiMo-V2.6-Pro core | No Shared Experts core— |
| Qwen3.5-397B-A17B core | 10 Routed + 1 Shared Experts Activated per Token core8 Routed + 1 Shared Experts Activated per Token core— |
| Qwen3.8-Flash-Next core | 10 Routed + 1 Shared Experts Activated per Token core— |