specific method · filed under model architecture
No Shared Experts
A mixture-of-experts configuration that has no shared experts.
Also called contains no shared experts, Sparse MoE FFNs without shared experts, without shared experts.
- sources
- 3
- models
- 3
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 2not used 1
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
both use sparse MoE FFNs without shared experts.
coremodel architecturein MiMo-V2.6-Pro-RLXiaomi
both use sparse MoE FFNs without shared experts.
coremodel architecturein MiMo-V2.6-Flash-RLXiaomi
and contains no shared experts.
not usedmodel architecturein MiMo-V2-FlashXiaomi
Filed alongside
Other methods under model architecture :: channel mixer :: mixture of experts :: shared experts.