ambiguous · filed under model architecture
KDA
A method labeled KDA and described as hybrid linear attention that replaces standard quadratic attention in some layers; the evidence does not settle whether this is Kimi Delta Attention or a different expansion of the acronym.
Also called Kimi Distributed Attention (KDA).
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
A hybrid linear attention mechanism that replaces standard quadratic attention in a subset of layers. KDA maintains full expressiveness for critical layers
Its architecture uses KDA and Attention Residuals for computational efficiency.
Filed alongside
Other methods under model architecture :: token mixer :: linear attention & state space :: gated delta network.