Model techniques map
Techniquesmodel architecturetoken mixerlinear attention & state spacegated delta network

ambiguous · filed under model architecture

KDA

A method labeled KDA and described as hybrid linear attention that replaces standard quadratic attention in some layers; the evidence does not settle whether this is Kimi Delta Attention or a different expansion of the acronym.

Also called Kimi Distributed Attention (KDA).

sources
2
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 2

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

A hybrid linear attention mechanism that replaces standard quadratic attention in a subset of layers. KDA maintains full expressiveness for critical layers

coremodel architecturein Kimi K3Moonshot AI

Its architecture uses KDA and Attention Residuals for computational efficiency.

coremodel architecturein Kimi K3Moonshot AI

Filed alongside

Other methods under model architecture :: token mixer :: linear attention & state space :: gated delta network.