taxonomy node · level 4
multi-head latent attention
4 methods filed at this node or below it, from the sources of 4 models.
model architecture :: token mixer :: softmax attention :: multi-head latent attention
Matching aids for the classifier: MLA; latent attention; gated MLA.
In this branch 4
Everything filed at this node or below it, with one collapsible heading per child node.
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| GLM-5.2 core | Multi-Head Latent Attention core— |
| MiniMax-M3 not used | Multi-Head Latent Attention not used— |
| DeepSeek-V3.2 core | MQA Mode of Multi-Head Latent Attention coreMulti-Head Latent Attention MHA/MQA Modes used— |
| Kimi K3 core | Gated Multi-Head Latent Attention core— |