implementation detail · filed under model architecture
MoE with routed and shared experts
An MoE layer configuration that includes both routed experts and a shared expert.
Also called 256 routed experts and 1 shared expert, Routed experts with shared expert.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
each containing 256 routed experts and 1 shared expert
coremodel architecturein Hy4-previewTencent Hunyuan
256 routed experts (top-8) and 1 shared expert
coremodel architecturein Hy4-previewTencent Hunyuan
Filed alongside
Other methods under model architecture :: channel mixer :: mixture of experts.
Mixture of ExpertsSparse expert activationDeepSeekMoEGated DeltaNet MoEMixture-of-Experts layersAsymmetric input/output activation splitDispatch recomputationFrequency-based partitioning of N-gram embedding slotsHybrid Mixture of ExpertsMegaMoEMoE with 128 routed experts and a shared expertMoE with 256 experts and top-8 routingMoE with 384 routed experts and a shared expertOutput-independent MoE gradient reformulationRouted expert output modulationSigmoid-based MoE router with auxiliary-loss-free load balancing