taxonomy node · level 3
linear attention & state space
16 methods filed at this node or below it, from the sources of 9 models.
model architecture :: token mixer :: linear attention & state space
Matching aids for the classifier: SSM; linear attention; lightning attention.
In this branch 16
Everything filed at this node or below it, with one collapsible heading per child node.
filed here 2
gated delta network 11
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | filed here | Mamba | gated delta network |
|---|---|---|---|
| GLM-5.3-Flash core | — | — | Kimi Delta Attention core— |
| NVIDIA-Nemotron-3-Ultra-550B-A55B core | — | Mamba coreMamba-2 coreMamba-2 SSM cache optional— | — |
| GLM-5.2 evaluated | — | — | Gated DeltaNet evaluatedSimpleGDN evaluated— |
| MiniMax-M3 mentioned | Linear attention mentionedLightning Attention not used— | — | — |
| Kimi K3 core | — | — | Hybrid linear attention coreKDA coreKimi Delta Attention coreKimi Delta Attention and Attention Residuals architecture coreGated DeltaNet usedKimi Delta Attention with input-dependent full-rank output gate usedKimi Delta Attention with lower-bounded log-decay used— |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B core | — | Mamba-2 core— | — |
| Qwen3.5-397B-A17B core | Linear attention core— | — | Gated DeltaNet coreGated DeltaNet with reduced KV-head configuration coreGated DeltaNet–sparse MoE hybrid core— |
| Qwen3.6-35B-A3B core | — | — | Gated DeltaNet core— |
| Qwen3.8-Flash-Next core | — | — | Gated DeltaNet coreGated DeltaNet with bounded sigmoid output gate used— |