Model techniques map
Techniquesmodel architecturetoken mixersparse attention

general family · filed under model architecture

Sparse attention

A broad attention family in which attention is restricted to a sparse subset rather than computed densely.

Also called sparse attention layers.

sources
6
models
2
labs adopt it
2
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

mentioned 1used 1core 4

Documented in

Further reading

Picked by hand, not extracted: where to read more, not evidence for anything on this page.

Evidence

6 spans quoted from the sources, strongest treatment first.

MSA is a clean and easily extensible new sparse attention architecture

coremodel architecturein MiniMax-M3MiniMax

uses sparse attention from the first step

coremodel architecturein MiniMax Sparse AttentionMiniMax

MiniMax Goes Sparse

coremodel architecturein MiniMax-M3MiniMax

Context Scaling via Sparse Attention: M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency.

coremodel architecturein MiniMax-M3MiniMax

every four sparse attention layers

usedmodel architecturein GLM-5.2Z.ai

Context Scaling via Sparse Attention

mentionedmodel architecturein MiniMax-M3MiniMax

Filed alongside

Other methods under model architecture :: token mixer :: sparse attention.