implementation detail · filed under model architecture
Gated Attention and Sliding-Window Attention Configuration
A specific attention-head configuration combining gated-attention and sliding-window-attention heads with a dense-attention setting.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
+ 48 GA / 64 SWA Q-heads, 𝑘dense=1 (final ablation architecture)
usedmodel architecturein Laguna XS.2Poolside
Filed alongside
Other methods under model architecture :: token mixer :: hybrid layer stacking.
Hybrid AttentionMamba-2, Mixture-of-Experts, and Selective Attention HybridHybrid Mamba-TransformerHybrid Mamba-AttentionGated DeltaNet and Gated AttentionHybrid Mamba-Transformer Mixture-of-Experts Layer LayoutDense Attention FallbackGated DeltaNet and Full AttentionGated DeltaNet and Qwen Sparse Attentionhybrid Gated DeltaNet + sparse MoE architectureHybrid sparse mixture-of-experts Transformer architectureSearch-Based Sliding-Window Attention Pattern