Model techniques map
Taxonomymodel architecturetoken mixersparse attention

taxonomy node · level 3

sparse attention

57 methods filed at this node or below it, from the sources of 11 models.

model architecture :: token mixer :: sparse attention

Matching aids for the classifier: sparse softmax attention; natively trained sparsity; two-stage sparse attention.

In this branch 57

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 24

DeepSeek Sparse Attention core · 15 sources · 17 quotes
Compressed Sparse Attention core · 6 sources · 6 quotes
Qwen Sparse Attention core · 6 sources · 6 quotes
Sparse attention core · 6 sources · 6 quotes
Heavily Compressed Attention core · 5 sources · 5 quotes
Gated DeepSeek Sparse Attention core · 3 sources · 3 quotes
Fixed-budget sparse-attention selection core · 1 source · 1 quote
NoPE sparse multi-head latent attention core · 1 source · 1 quote
QSA micro-block compression at ratio 4 core · 1 source · 1 quote
Sequential block processing core · 1 source · 1 quote
Token-wise compression core · 1 source · 1 quote
KV-outer sparse attention used · 2 sources · 2 quotes
CSA2 Full Mode used · 1 source · 2 quotes
Natively trained sparsity used · 1 source · 1 quote
Sparse softmax attention used · 1 source · 1 quote
Two-stage introduction of sparse attention used · 1 source · 1 quote
Two-stage sparse attention used · 1 source · 1 quote
Sparse retrieval over long contexts evaluated · 1 source · 1 quote
Native Sparse Attention not used · 1 source · 1 quote

block selection 9

MiniMax Sparse Attention core · 6 sources · 6 quotes
Block-level selection core · 1 source · 1 quote
Dynamic sparse selection core · 1 source · 1 quote
Main Branch core · 1 source · 2 quotes
Top-512 block selection core · 1 source · 1 quote
Top-k block selection used · 2 sources · 2 quotes
Block Max Pooling used · 1 source · 1 quote
Top-scoring compressed-block selection used · 1 source · 1 quote
Mixture of Block Attention evaluated · 1 source · 1 quote

sparse attention indexer 22

IndexCache core · 3 sources · 3 quotes
IndexShare core · 3 sources · 3 quotes
Lightning Indexer core · 3 sources · 3 quotes
Compressed Sparse Attention 2 core · 2 sources · 6 quotes
Fine-grained token selection core · 2 sources · 2 quotes
Hierarchical Sparse Indexer core · 2 sources · 5 quotes
Compressed lightweight indexer core · 1 source · 1 quote
Index Branch core · 1 source · 2 quotes
IndexPool core · 1 source · 1 quote
MQA indexer core · 1 source · 1 quote
Reindex Mode core · 1 source · 3 quotes
Reuse Mode core · 1 source · 3 quotes
Dense Warm-up Stage used · 2 sources · 2 quotes
Average pooling used · 1 source · 1 quote
Block-causal scoring used · 1 source · 1 quote
Detached indexer-input optimization used · 1 source · 1 quote
Indexer Warmup used · 1 source · 2 quotes
Single-head index key used · 1 source · 1 quote
Index Branch output not used · 1 source · 1 quote
Index Branch value head not used · 1 source · 1 quote

fixed-pattern sparse attention 2

Sparse Global Attention Anchors core · 1 source · 1 quote
Local Block used · 1 source · 2 quotes

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modelfiled hereblock selectionsparse attention indexerfixed-pattern sparse attention
GLM-5.3-Flash coreNoPE sparse multi-head latent attention core——IndexPool core——
DeepSeek-V4.1-Flash coreCross-layer KV and index reuse with statically assigned CSA2 modes coreCSA2 Full Mode usedFrom-scratch sparse attention training without dense warmup usedProgressive sequence-length extension for sparse attention usedSparse retrieval over long contexts evaluated——Compressed Sparse Attention 2 coreHierarchical Sparse Indexer coreReindex Mode coreReuse Mode coreCross-stage shared-state management for attention reuse used——
Hy4-preview coreGated DeepSeek Sparse Attention core——IndexCache core——
NVIDIA-Nemotron-3-Ultra-550B-A55B core———Sparse Global Attention Anchors core—
DeepSeek-V4-Flash coreCompressed Sparse Attention coreDeepSeek Sparse Attention coreHeavily Compressed Attention coreToken-wise compression coreTwo-stage introduction of sparse attention used————
GLM-5.2 coreDeepSeek Sparse Attention coreSparse attention used——IndexShare coreLightning Indexer used——
MiniMax-M3 coreSequential block processing coreSparse attention coreKV-outer sparse attention usedNatively trained sparsity usedSparse softmax attention usedTwo-stage sparse attention usedDeepSeek Sparse Attention evaluatedCompressed Sparse Attention not usedNative Sparse Attention not used—Block-level selection coreDynamic sparse selection coreMain Branch coreMiniMax Sparse Attention coreBlock Max Pooling usedTop-k block selection usedMixture of Block Attention evaluated—Index Branch coreIndexer Warmup usedSingle-head index key usedIndex Branch output not usedIndex Branch value head not used—Local Block used—
DeepSeek-V3.2 coreDeepSeek Sparse Attention coreSparse-attention continued pre-training with joint model and indexer optimization used——Fine-grained token selection coreLightning Indexer coreDense Warm-up Stage usedDetached indexer-input optimization used——
DeepSeek-V4-Pro coreCompressed Sparse Attention coreDeepSeek Sparse Attention coreHeavily Compressed Attention coreToken-wise compression coreTwo-stage introduction of sparse attention used————
Qwen3.5-397B-A17B coreDeepSeek Sparse Attention core————
Qwen3.8-Flash-Next coreFixed-budget sparse-attention selection coreQSA micro-block compression at ratio 4 coreQwen Sparse Attention coreJoint backbone and indexer training under sparse attention used—Top-512 block selection coreTop-scoring compressed-block selection used—Compressed lightweight indexer coreMQA indexer coreReuse QSA index selection across speculative decoding steps coreAverage pooling usedBlock-causal scoring used——