general family · filed under model architecture
Hybrid sparse mixture-of-experts Transformer architecture
Also called hybrid sparse MoE Transformer backbone.
- source
- 1
- models
- 2
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
MiMo-V2.6 builds a hybrid sparse MoE Transformer backbone that interleaves Local Sliding Window Attention (SWA) with Global Attention (GA)
coremodel architecturein MiMo-V2.6Xiaomi
Filed alongside
Other methods under model architecture :: token mixer :: hybrid layer stacking.
Hybrid AttentionMamba-2, Mixture-of-Experts, and Selective Attention HybridHybrid Mamba-TransformerHybrid Mamba-AttentionGated DeltaNet and Gated AttentionHybrid Mamba-Transformer Mixture-of-Experts Layer LayoutDense Attention FallbackGated Attention and Sliding-Window Attention ConfigurationGated DeltaNet and Full AttentionGated DeltaNet and Qwen Sparse Attentionhybrid Gated DeltaNet + sparse MoE architectureSearch-Based Sliding-Window Attention Pattern