Model techniques map
Techniquesmodel architecturetoken mixerhybrid layer stacking

general family · filed under model architecture

Hybrid sparse mixture-of-experts Transformer architecture

Also called hybrid sparse MoE Transformer backbone.

source
1
models
2
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

MiMo-V2.6 builds a hybrid sparse MoE Transformer backbone that interleaves Local Sliding Window Attention (SWA) with Global Attention (GA)

coremodel architecturein MiMo-V2.6Xiaomi

Filed alongside

Other methods under model architecture :: token mixer :: hybrid layer stacking.