taxonomy node · level 4
attention sink
3 methods filed at this node or below it, from the sources of 5 models.
model architecture :: token mixer :: softmax attention :: attention sink
Matching aids for the classifier: learnable attention sink bias; forced sink; RoPE with attention sink.
In this branch 3
Everything filed at this node or below it, with one collapsible heading per child node.
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| DeepSeek-V4-Flash used | Attention sink used— |
| MiMo-V2.5 used | Learnable attention sink bias used— |
| MiniMax-M3 not used | Learnable attention sink bias not usedRoPE with attention sink unclear— |
| DeepSeek-V4-Pro used | Attention sink used— |
| MiMo-V2.5-Pro used | Learnable attention sink bias used— |