Model techniques map
Taxonomymodel architecturetoken mixersoftmax attention

taxonomy node · level 3

softmax attention

32 methods filed at this node or below it, from the sources of 20 models.

model architecture :: token mixer :: softmax attention

Matching aids for the classifier: dense attention; full attention; attention layers.

In this branch 32

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 5

Attention layers core · 1 source · 1 quote
DFlash attention core · 1 source · 1 quote
Softplus attention gating core · 1 source · 1 quote
Learned softmax denominator bias used · 1 source · 1 quote
Cascade attention not used · 2 sources · 2 quotes

global attention 4

Global attention core · 4 sources · 4 quotes
Unified Keys and Values core · 2 sources · 2 quotes
Full attention used · 2 sources · 2 quotes
Multi-Head Attention (MHA) used · 1 source · 1 quote

sliding window attention 12

Sliding Window Attention core · 8 sources · 10 quotes
Hybrid Sliding Window Attention core · 4 sources · 5 quotes
Sink-Augmented Sliding Window Attention core · 1 source · 1 quote
512-token Sliding Window Attention used · 2 sources · 2 quotes
Decoder SWA Bounded Replay used · 1 source · 2 quotes
Dynamic Attention Window Size Training used · 1 source · 1 quote
Pure Sliding Window Attention used · 1 source · 1 quote
SWA-128 used · 1 source · 1 quote
Fixed Sliding Window Attention evaluated · 1 source · 1 quote
Per-Head Gating evaluated · 1 source · 1 quote
SWA-1024 evaluated · 1 source · 1 quote

grouped-query attention 4

Grouped-query attention core · 11 sources · 13 quotes
Block-sparse grouped-query attention core · 1 source · 1 quote
Multi-query attention used · 2 sources · 2 quotes
Softplus-based per-head gating used · 1 source · 1 quote

multi-head latent attention 4

Gated Multi-Head Latent Attention core · 3 sources · 3 quotes
MQA Mode of Multi-Head Latent Attention core · 2 sources · 2 quotes
Multi-Head Latent Attention core · 2 sources · 3 quotes
Multi-Head Latent Attention MHA/MQA Modes used · 1 source · 1 quote

attention sink 3

Learnable attention sink bias used · 5 sources · 5 quotes
Attention sink used · 1 source · 1 quote
RoPE with attention sink unclear · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modelfiled hereglobal attentionsliding window attentiongrouped-query attentionmulti-head latent attentionattention sink
DeepSeek-V4.1-Flash core——Sliding Window Attention coreDecoder SWA Bounded Replay usedPure Sliding Window Attention used————
NVIDIA-Nemotron-3-Ultra-550B-A55B coreAttention layers core——————
MiMo-V2.6-Flash core——Alternating Row-Major and Column-Major Token Serialization coreSink-Augmented Sliding Window Attention core————
DeepSeek-V4-Flash used——Sliding Window Attention used—Multi-query attention used——Attention sink used—
MiMo-V2.5 core—Global attention coreFull attention used—Hybrid Sliding Window Attention coreSliding Window Attention coreSWA-128 used—Grouped-query attention used——Learnable attention sink bias used—
Hy3 core———Grouped-query attention core———
GLM-5.2 core——Sliding Window Attention evaluated—Multi-query attention mentioned—Multi-Head Latent Attention core——
MiniMax-M3 core—Full attention evaluated—Fixed Sliding Window Attention evaluatedSliding Window Attention mentioned—Block-sparse grouped-query attention coreGrouped-query attention core—Multi-Head Latent Attention not used—Learnable attention sink bias not usedRoPE with attention sink unclear—
DeepSeek-V3.2 core————MQA Mode of Multi-Head Latent Attention coreMulti-Head Latent Attention MHA/MQA Modes used——
DeepSeek-V4-Flash-Vision-Exp coreDFlash attention core——————
DeepSeek-V4-Pro used——Sliding Window Attention used—Multi-query attention used——Attention sink used—
Gemma 4 31B core—Unified Keys and Values core—————
Inkling used——512-token Sliding Window Attention used—Grouped-query attention used———
Kimi K3 core————Gated Multi-Head Latent Attention core——
Laguna-S-2.1 coreSoftplus attention gating core——512-token Sliding Window Attention evaluatedPer-Head Gating evaluatedSWA-1024 evaluated—Grouped-query attention coreSoftplus-based per-head gating used———
MiMo-V2.5-Pro used—Global attention used—Hybrid Sliding Window Attention usedSliding Window Attention used———Learnable attention sink bias used—
MiMo-V2.6-Pro core——Alternating Row-Major and Column-Major Token Serialization coreSink-Augmented Sliding Window Attention core————
Qwen3.5-397B-A17B used—Multi-Head Attention (MHA) used—Dynamic Attention Window Size Training used————
Step-3.7-Flash not usedCascade attention not used——————
gpt-oss-120b usedLearned softmax denominator bias used———Grouped-query attention used———

Proposed children

Paths the corpus wanted and the taxonomy does not have. A human promotes them into the outline; the extraction cannot.

gated attentionGated Attention, attention gating