specific method · filed under model architecture
Gated DeepSeek Sparse Attention
A gated variant of DeepSeek Sparse Attention; the evidence also mentions its use with IndexCache.
- sources
- 3
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 3
Documented in
Further reading
Picked by hand, not extracted: where to read more, not evidence for anything on this page.
- DeepSeek Sparse Attention - LLM Architecture Gallery explainer sebastianraschka.combackground on the DSA mechanism being gated
- Gated Attention - LLM Architecture Gallery explainer sebastianraschka.combackground on the gating mechanism
Evidence
3 spans quoted from the sources, strongest treatment first.
the attention module employs Gated DeepSeek Sparse Attention (Gated DSA)
coremodel architecturein Hy4-previewTencent Hunyuan
the attention module employs Gated DeepSeek Sparse Attention (Gated DSA) with IndexCache for cross-layer sparse index reuse
coremodel architecturein Hy4-previewTencent Hunyuan
the attention module employs Gated DeepSeek Sparse Attention (Gated DSA) with IndexCache
coremodel architecturein Hy4-previewTencent Hunyuan
Filed alongside
Other methods under model architecture :: token mixer :: sparse attention.
DeepSeek Sparse AttentionCompressed Sparse AttentionQwen Sparse AttentionSparse attentionHeavily Compressed AttentionCSA2 Full ModeKV-outer sparse attentionProgressive sequence-length extension for sparse attentionSparse-attention continued pre-training with joint model and indexer optimizationCross-layer KV and index reuse with statically assigned CSA2 modesFixed-budget sparse-attention selectionFrom-scratch sparse attention training without dense warmupJoint backbone and indexer training under sparse attentionNative Sparse AttentionNatively trained sparsityNoPE sparse multi-head latent attentionQSA micro-block compression at ratio 4Sequential block processingSparse retrieval over long contextsSparse softmax attentionToken-wise compressionTwo-stage introduction of sparse attentionTwo-stage sparse attention