Model techniques map
Taxonomymodel architecturetoken mixer

taxonomy node · level 2

token mixer

123 methods filed at this node or below it, from the sources of 25 models.

model architecture :: token mixer

Matching aids for the classifier: sequence mixing; attention block.

In this branch 123

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 5

softmax attention 32

Attention layers core · 1 source · 1 quote
DFlash attention core · 1 source · 1 quote
Softplus attention gating core · 1 source · 1 quote
Learned softmax denominator bias used · 1 source · 1 quote
Cascade attention not used · 2 sources · 2 quotes

global attention 4

Global attention core · 4 sources · 4 quotes
Unified Keys and Values core · 2 sources · 2 quotes
Full attention used · 2 sources · 2 quotes
Multi-Head Attention (MHA) used · 1 source · 1 quote

sliding window attention 12

Sliding Window Attention core · 8 sources · 10 quotes
Hybrid Sliding Window Attention core · 4 sources · 5 quotes
Sink-Augmented Sliding Window Attention core · 1 source · 1 quote
512-token Sliding Window Attention used · 2 sources · 2 quotes
Decoder SWA Bounded Replay used · 1 source · 2 quotes
Dynamic Attention Window Size Training used · 1 source · 1 quote
Pure Sliding Window Attention used · 1 source · 1 quote
SWA-128 used · 1 source · 1 quote
Fixed Sliding Window Attention evaluated · 1 source · 1 quote
Per-Head Gating evaluated · 1 source · 1 quote
SWA-1024 evaluated · 1 source · 1 quote

grouped-query attention 4

Grouped-query attention core · 11 sources · 13 quotes
Block-sparse grouped-query attention core · 1 source · 1 quote
Multi-query attention used · 2 sources · 2 quotes
Softplus-based per-head gating used · 1 source · 1 quote

multi-head latent attention 4

Gated Multi-Head Latent Attention core · 3 sources · 3 quotes
MQA Mode of Multi-Head Latent Attention core · 2 sources · 2 quotes
Multi-Head Latent Attention core · 2 sources · 3 quotes
Multi-Head Latent Attention MHA/MQA Modes used · 1 source · 1 quote

attention sink 3

Learnable attention sink bias used · 5 sources · 5 quotes
Attention sink used · 1 source · 1 quote
RoPE with attention sink unclear · 1 source · 1 quote

sparse attention 57

DeepSeek Sparse Attention core · 15 sources · 17 quotes
Compressed Sparse Attention core · 6 sources · 6 quotes
Qwen Sparse Attention core · 6 sources · 6 quotes
Sparse attention core · 6 sources · 6 quotes
Heavily Compressed Attention core · 5 sources · 5 quotes
Gated DeepSeek Sparse Attention core · 3 sources · 3 quotes
Fixed-budget sparse-attention selection core · 1 source · 1 quote
NoPE sparse multi-head latent attention core · 1 source · 1 quote
QSA micro-block compression at ratio 4 core · 1 source · 1 quote
Sequential block processing core · 1 source · 1 quote
Token-wise compression core · 1 source · 1 quote
KV-outer sparse attention used · 2 sources · 2 quotes
CSA2 Full Mode used · 1 source · 2 quotes
Natively trained sparsity used · 1 source · 1 quote
Sparse softmax attention used · 1 source · 1 quote
Two-stage introduction of sparse attention used · 1 source · 1 quote
Two-stage sparse attention used · 1 source · 1 quote
Sparse retrieval over long contexts evaluated · 1 source · 1 quote
Native Sparse Attention not used · 1 source · 1 quote

block selection 9

MiniMax Sparse Attention core · 6 sources · 6 quotes
Block-level selection core · 1 source · 1 quote
Dynamic sparse selection core · 1 source · 1 quote
Main Branch core · 1 source · 2 quotes
Top-512 block selection core · 1 source · 1 quote
Top-k block selection used · 2 sources · 2 quotes
Block Max Pooling used · 1 source · 1 quote
Top-scoring compressed-block selection used · 1 source · 1 quote
Mixture of Block Attention evaluated · 1 source · 1 quote

sparse attention indexer 22

IndexCache core · 3 sources · 3 quotes
IndexShare core · 3 sources · 3 quotes
Lightning Indexer core · 3 sources · 3 quotes
Compressed Sparse Attention 2 core · 2 sources · 6 quotes
Fine-grained token selection core · 2 sources · 2 quotes
Hierarchical Sparse Indexer core · 2 sources · 5 quotes
Compressed lightweight indexer core · 1 source · 1 quote
Index Branch core · 1 source · 2 quotes
IndexPool core · 1 source · 1 quote
MQA indexer core · 1 source · 1 quote
Reindex Mode core · 1 source · 3 quotes
Reuse Mode core · 1 source · 3 quotes
Dense Warm-up Stage used · 2 sources · 2 quotes
Average pooling used · 1 source · 1 quote
Block-causal scoring used · 1 source · 1 quote
Detached indexer-input optimization used · 1 source · 1 quote
Indexer Warmup used · 1 source · 2 quotes
Single-head index key used · 1 source · 1 quote
Index Branch output not used · 1 source · 1 quote
Index Branch value head not used · 1 source · 1 quote

fixed-pattern sparse attention 2

Sparse Global Attention Anchors core · 1 source · 1 quote
Local Block used · 1 source · 2 quotes

linear attention & state space 16

Linear attention core · 2 sources · 2 quotes
Lightning Attention not used · 1 source · 1 quote

Mamba 3

Mamba-2 core · 5 sources · 6 quotes
Mamba core · 2 sources · 2 quotes
Mamba-2 SSM cache optional · 1 source · 1 quote

gated delta network 11

Gated DeltaNet core · 13 sources · 13 quotes
Kimi Delta Attention core · 6 sources · 6 quotes
Gated DeltaNet–sparse MoE hybrid core · 2 sources · 2 quotes
KDA core · 2 sources · 2 quotes
Hybrid linear attention core · 1 source · 1 quote
SimpleGDN evaluated · 1 source · 1 quote

hybrid layer stacking 13

Hybrid Attention core · 27 sources · 29 quotes
Hybrid Mamba-Transformer core · 4 sources · 6 quotes
Hybrid Mamba-Attention core · 3 sources · 4 quotes
Gated DeltaNet and Full Attention core · 1 source · 1 quote
Gated DeltaNet and Gated Attention core · 1 source · 2 quotes
Gated DeltaNet and Qwen Sparse Attention core · 1 source · 1 quote
Search-Based Sliding-Window Attention Pattern evaluated · 1 source · 1 quote
Dense Attention Fallback unclear · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modelfiled heresoftmax attentionsparse attentionlinear attention & state spacehybrid layer stacking
GLM-5.3-Flash core——IndexPool coreNoPE sparse multi-head latent attention core—Kimi Delta Attention core—Hybrid Attention core—
DeepSeek-V4.1-Flash coreCross-layer attention sharing coreMicro-batch-level shared-state lifetime management used—Sliding Window Attention coreDecoder SWA Bounded Replay usedPure Sliding Window Attention used—Compressed Sparse Attention 2 coreCross-layer KV and index reuse with statically assigned CSA2 modes coreHierarchical Sparse Indexer coreReindex Mode coreReuse Mode coreCross-stage shared-state management for attention reuse usedCSA2 Full Mode usedFrom-scratch sparse attention training without dense warmup usedProgressive sequence-length extension for sparse attention usedSparse retrieval over long contexts evaluated———
Hy4-preview core——Gated DeepSeek Sparse Attention coreIndexCache core———
NVIDIA-Nemotron-3-Ultra-550B-A55B core—Attention layers core—Sparse Global Attention Anchors core—Mamba coreMamba-2 coreMamba-2 SSM cache optional—Hybrid Mamba-Attention coreHybrid Mamba-Transformer core—
MiMo-V2.6-Flash core—Alternating Row-Major and Column-Major Token Serialization coreSink-Augmented Sliding Window Attention core———Hybrid Attention coreHybrid sparse mixture-of-experts Transformer architecture core—
DeepSeek-V4-Flash core—Attention sink usedMulti-query attention usedSliding Window Attention used—Compressed Sparse Attention coreDeepSeek Sparse Attention coreHeavily Compressed Attention coreToken-wise compression coreTwo-stage introduction of sparse attention used——Hybrid Attention used—
MiMo-V2.5 core—Global attention coreHybrid Sliding Window Attention coreSliding Window Attention coreFull attention usedGrouped-query attention usedLearnable attention sink bias usedSWA-128 used———Hybrid Attention core—
Hy3 core—Grouped-query attention core————
GLM-5.2 core—Multi-Head Latent Attention coreSliding Window Attention evaluatedMulti-query attention mentioned—DeepSeek Sparse Attention coreIndexShare coreLightning Indexer usedSparse attention used—Gated DeltaNet evaluatedSimpleGDN evaluated—Search-Based Sliding-Window Attention Pattern evaluated—
MiniMax-M3 core—Block-sparse grouped-query attention coreGrouped-query attention coreFixed Sliding Window Attention evaluatedFull attention evaluatedSliding Window Attention mentionedLearnable attention sink bias not usedMulti-Head Latent Attention not usedRoPE with attention sink unclear—Block-level selection coreDynamic sparse selection coreIndex Branch coreMain Branch coreMiniMax Sparse Attention coreSequential block processing coreSparse attention coreBlock Max Pooling usedIndexer Warmup usedKV-outer sparse attention usedLocal Block usedNatively trained sparsity usedSingle-head index key usedSparse softmax attention usedTop-k block selection usedTwo-stage sparse attention usedDeepSeek Sparse Attention evaluatedMixture of Block Attention evaluatedCompressed Sparse Attention not usedIndex Branch output not usedIndex Branch value head not usedNative Sparse Attention not used—Linear attention mentionedLightning Attention not used—Dense Attention Fallback unclear—
DeepSeek-V3.2 core—MQA Mode of Multi-Head Latent Attention coreMulti-Head Latent Attention MHA/MQA Modes used—DeepSeek Sparse Attention coreFine-grained token selection coreLightning Indexer coreDense Warm-up Stage usedDetached indexer-input optimization usedSparse-attention continued pre-training with joint model and indexer optimization used———
DeepSeek-V4-Flash-Vision-Exp core—DFlash attention core————
DeepSeek-V4-Pro core—Attention sink usedMulti-query attention usedSliding Window Attention used—Compressed Sparse Attention coreDeepSeek Sparse Attention coreHeavily Compressed Attention coreToken-wise compression coreTwo-stage introduction of sparse attention used——Hybrid Attention used—
Gemma 4 31B core—Unified Keys and Values core———Hybrid Attention core—
Inkling coreKernel-4 convolutional layers in decoder blocks used—512-token Sliding Window Attention usedGrouped-query attention used———Hybrid Attention core—
Kimi K3 coreFactorized spatial-temporal video attention with temporal pooling used—Gated Multi-Head Latent Attention core——Hybrid linear attention coreKDA coreKimi Delta Attention coreKimi Delta Attention and Attention Residuals architecture coreGated DeltaNet usedKimi Delta Attention with input-dependent full-rank output gate usedKimi Delta Attention with lower-bounded log-decay used—Hybrid Attention core—
Laguna-S-2.1 coreDense gated attention with full RoPE and full gating evaluated—Grouped-query attention coreSoftplus attention gating coreSoftplus-based per-head gating used512-token Sliding Window Attention evaluatedPer-Head Gating evaluatedSWA-1024 evaluated———Hybrid Attention coreGated Attention and Sliding-Window Attention Configuration used—
MiMo-V2.5-Pro core—Global attention usedHybrid Sliding Window Attention usedLearnable attention sink bias usedSliding Window Attention used———Hybrid Attention core—
MiMo-V2.6-Pro core—Alternating Row-Major and Column-Major Token Serialization coreSink-Augmented Sliding Window Attention core———Hybrid Attention coreHybrid sparse mixture-of-experts Transformer architecture core—
NVIDIA-Nemotron-3.5-Lightning-30B-A3B core———Mamba-2 core—Hybrid Mamba-Transformer coreMamba-2, Mixture-of-Experts, and Selective Attention Hybrid core—
Qwen3.5-397B-A17B core—Dynamic Attention Window Size Training usedMulti-Head Attention (MHA) used—DeepSeek Sparse Attention core—Gated DeltaNet coreGated DeltaNet with reduced KV-head configuration coreGated DeltaNet–sparse MoE hybrid coreLinear attention core—Gated DeltaNet and Full Attention corehybrid Gated DeltaNet + sparse MoE architecture coreHybrid Mamba-Transformer Mixture-of-Experts Layer Layout core—
Qwen3.6-35B-A3B core———Gated DeltaNet core—Gated DeltaNet and Gated Attention core—
Qwen3.8-Flash-Next core——Compressed lightweight indexer coreFixed-budget sparse-attention selection coreMQA indexer coreQSA micro-block compression at ratio 4 coreQwen Sparse Attention coreReuse QSA index selection across speculative decoding steps coreTop-512 block selection coreAverage pooling usedBlock-causal scoring usedJoint backbone and indexer training under sparse attention usedTop-scoring compressed-block selection used—Gated DeltaNet coreGated DeltaNet with bounded sigmoid output gate used—Gated DeltaNet and Qwen Sparse Attention coreHybrid Attention core—
Step-3.7-Flash not used—Cascade attention not used————
gpt-oss-120b used—Grouped-query attention usedLearned softmax denominator bias used———Hybrid Attention used—