taxonomy node · level 4
grouped-query attention
4 methods filed at this node or below it, from the sources of 9 models.
model architecture :: token mixer :: softmax attention :: grouped-query attention
Matching aids for the classifier: GQA; multi-query attention; MQA; block-sparse GQA.
In this branch 4
Everything filed at this node or below it, with one collapsible heading per child node.
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| DeepSeek-V4-Flash used | Multi-query attention used— |
| MiMo-V2.5 used | Grouped-query attention used— |
| Hy3 core | Grouped-query attention core— |
| GLM-5.2 mentioned | Multi-query attention mentioned— |
| MiniMax-M3 core | Block-sparse grouped-query attention coreGrouped-query attention core— |
| DeepSeek-V4-Pro used | Multi-query attention used— |
| Inkling used | Grouped-query attention used— |
| Laguna-S-2.1 core | Grouped-query attention coreSoftplus-based per-head gating used— |
| gpt-oss-120b used | Grouped-query attention used— |