specific method · filed under model architecture
Gated DeltaNet and Gated Attention
Combines Gated DeltaNet linear-attention layers with standard gated-attention layers.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
It uses a hybrid sparse mixture-of-experts architecture combining Gated DeltaNet linear attention with standard gated attention layers
coremodel architecturein Qwen3.6-35B-A3BAlibaba Cloud
combining Gated DeltaNet linear attention with standard gated attention layers
coremodel architecturein Qwen3.6-35B-A3BAlibaba Cloud
Filed alongside
Other methods under model architecture :: token mixer :: hybrid layer stacking.
Hybrid AttentionMamba-2, Mixture-of-Experts, and Selective Attention HybridHybrid Mamba-TransformerHybrid Mamba-AttentionHybrid Mamba-Transformer Mixture-of-Experts Layer LayoutDense Attention FallbackGated Attention and Sliding-Window Attention ConfigurationGated DeltaNet and Full AttentionGated DeltaNet and Qwen Sparse Attentionhybrid Gated DeltaNet + sparse MoE architectureHybrid sparse mixture-of-experts Transformer architectureSearch-Based Sliding-Window Attention Pattern