Model techniques map
Taxonomymodel architecturechannel mixer

taxonomy node · level 2

channel mixer

68 methods filed at this node or below it, from the sources of 25 models.

model architecture :: channel mixer

Matching aids for the classifier: feed-forward block; FFN.

In this branch 68

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 1

dense feed-forward network 6

SiTU-GLU core · 2 sources · 3 quotes
SwiGLU used · 3 sources · 3 quotes
SwiGLU clamping used · 3 sources · 3 quotes
Dense feed-forward network used · 1 source · 1 quote
SiTU used · 1 source · 1 quote
SwiGLU clipping not used · 1 source · 1 quote

mixture of experts 61

Mixture of Experts core · 88 sources · 101 quotes
Sparse expert activation core · 8 sources · 8 quotes
DeepSeekMoE core · 2 sources · 2 quotes
Mixture-of-Experts layers core · 2 sources · 2 quotes
MoE with routed and shared experts core · 2 sources · 2 quotes
Asymmetric input/output activation split core · 1 source · 1 quote
Dispatch recomputation core · 1 source · 1 quote
Gated DeltaNet MoE core · 1 source · 2 quotes
Hybrid Mixture of Experts core · 1 source · 1 quote
MegaMoE used · 1 source · 1 quote
MoE with 256 experts and top-8 routing used · 1 source · 1 quote
Routed expert output modulation used · 1 source · 1 quote

expert routing 15

Top-8 expert routing core · 5 sources · 5 quotes
Token-level expert routing core · 3 sources · 3 quotes
Keep Routing core · 1 source · 1 quote
Routed experts core · 1 source · 1 quote
Token-choice routing with softplus gating core · 1 source · 1 quote
Top-6 expert routing core · 1 source · 1 quote
Fixed Top-k routing with frozen bias default · 1 source · 1 quote
Anticipatory Routing used · 2 sources · 2 quotes
DP-aware routing used · 1 source · 1 quote
Hash routing used · 1 source · 1 quote
Latent-space routing used · 1 source · 1 quote
Token-choice routing used · 1 source · 1 quote
Top-4 expert routing used · 1 source · 2 quotes
Loss-spike-triggered Anticipatory Routing optional · 1 source · 1 quote

expert load balancing 15

Auxiliary-loss-free load balancing core · 3 sources · 5 quotes
Quantile Balancing core · 3 sources · 5 quotes
Auxiliary-loss load balancing used · 1 source · 1 quote
Exact coordinate minimization used · 1 source · 1 quote
Expert bias update factor used · 1 source · 1 quote
Expert Parallelism Load Balancing used · 1 source · 1 quote
formHC used · 1 source · 1 quote
Histogram-based quantile estimation used · 1 source · 2 quotes
Load balancing used · 1 source · 1 quote
Persistent load balancing used · 1 source · 1 quote
Pooled-global-batch quantile estimation used · 1 source · 1 quote
Round-robin load balancing optional · 1 source · 1 quote
MaxVio evaluated · 1 source · 1 quote

shared experts 7

Shared Experts core · 7 sources · 7 quotes
No Shared Experts core · 3 sources · 3 quotes
Shared Experts Active on Every Token core · 2 sources · 2 quotes
MoE with Expert Sinks core · 1 source · 1 quote

fine-grained experts 2

Granular Mixture of Experts not used · 2 sources · 2 quotes

latent mixture of experts 5

LatentMoE core · 7 sources · 9 quotes
Normalized LatentMoE core · 6 sources · 7 quotes
Hybrid Latent Mixture of Experts core · 1 source · 1 quote
Latent-dimension reduction used · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modelfiled heredense feed-forward networkmixture of experts
GLM-5.3-Flash core——Mixture of Experts coreToken-level expert routing core—
DeepSeek-V4.1-Flash core—SwiGLU usedSwiGLU clamping used—Asymmetric input/output activation split coreAuxiliary-loss-free load balancing coreDeepSeekMoE shared and fine-grained routed experts coreMixture of Experts coreModality-specific auxiliary-loss-free load balancing coreMoE with 384 routed experts and a shared expert coreformHC usedMixture-of-Experts layers usedShared Experts used—
Hy4-preview core——Mixture of Experts coreMoE with routed and shared experts coreShared Experts coreTop-8 expert routing coreTop-8 Routed-Expert Selection with Shared-Expert Activation core—
NVIDIA-Nemotron-3-Ultra-550B-A55B core——Hybrid Latent Mixture of Experts coreLatentMoE coreMixture of Experts coreMixture-of-Experts layers coreExpert Parallelism Load Balancing usedShared Experts usedMaxVio evaluatedGranular Mixture of Experts not used—
MiMo-V2.6-Flash core——Mixture of Experts coreNo Shared Experts core—
DeepSeek-V4-Flash core—SwiGLU usedSwiGLU clamping used—DeepSeekMoE coreMixture of Experts coreShared Experts coreAuxiliary-loss-free load balancing usedHash routing usedMegaMoE usedAnticipatory Routing optionalLoss-spike-triggered Anticipatory Routing optional—
MiMo-V2.5 core—Dense feed-forward network used—Mixture of Experts coreExpert bias update factor usedLoad balancing usedMoE with 256 experts and top-8 routing usedTop-8 expert routing usedRound-robin load balancing optionalNo Shared Experts not used—
Hy3 core——Mixture of Experts coreRouted experts coreShared Experts coreTop-8 expert routing core—
GLM-5.2 core——Mixture of Experts coreDP-aware routing used—
MiniMax-M3 core——Mixture of Experts corePersistent load balancing usedShared Experts usedTop-4 expert routing used—
DeepSeek-V3.2 core——Keep Routing core—
DeepSeek-V4-Flash-Vision-Exp core——Mixture of Experts core—
DeepSeek-V4-Pro core—SwiGLU usedSwiGLU clamping used—DeepSeekMoE coreMixture of Experts coreShared Experts coreAnticipatory Routing usedAuxiliary-loss-free load balancing usedHash routing usedMegaMoE usedLoss-spike-triggered Anticipatory Routing optional—
Gemma 4 31B core——Mixture of Experts coreSparse expert activation core—
Inkling core——Mixture of Experts coreMoE with Expert Sinks coreShared Experts Active on Every Token coreToken-level expert routing coreTop-6 expert routing coreSigmoid-based MoE router with auxiliary-loss-free load balancing used—
Kimi K3 core—SiTU-GLU coreSiTU used—Auxiliary-loss-free load balancing coreDispatch recomputation coreLatentMoE coreMixture of Experts coreNormalized LatentMoE coreOutput-independent MoE gradient reformulation coreQuantile Balancing coreSharded latent weights with fused all-gather GEMM epilogue coreSparse expert activation coreFixed Top-k routing with frozen bias defaultExact coordinate minimization usedExponential moving average of estimated quantiles usedHistogram-based quantile estimation usedLatent-space routing usedPooled-global-batch quantile estimation used—
Laguna-S-2.1 core——Mixture of Experts coreSparse expert activation coreToken-choice routing with softplus gating coreAuxiliary-loss load balancing usedRouted expert output modulation usedShared Experts usedToken-choice routing used—
MiMo-V2.5-Pro core——Mixture of Experts core—
MiMo-V2.6-Pro core——Mixture of Experts coreNo Shared Experts coreSparse expert activation core—
NVIDIA-Nemotron-3.5-Lightning-30B-A3B core——LatentMoE coreMixture of Experts coreToken-level expert routing coreLatent-dimension reduction usedMoE with 128 routed experts and a shared expert used—
Qwen3.5-397B-A17B core——10 Routed + 1 Shared Experts Activated per Token core8 Routed + 1 Shared Experts Activated per Token coreHybrid Mixture of Experts coreMixture of Experts coreSparse expert activation core—
Qwen3.6-35B-A3B core——Gated DeltaNet MoE coreMixture of Experts core—
Qwen3.8-Flash-Next coreContextual gating for N-gram embedding injection used—SwiGLU clipping not used—10 Routed + 1 Shared Experts Activated per Token coreMixture of Experts coreFrequency-based partitioning of N-gram embedding slots evaluated—
Step-3.7-Flash core——Mixture of Experts core—
gpt-oss-120b core—SwiGLU used—Mixture of Experts coreTop-k expert routing with softmax over selected experts core—