implementation detail · filed under model architecture
Gated DeltaNet with reduced KV-head configuration
A Gated Attention configuration with 32 query heads and 2 key/value heads; the evidence does not establish a distinct named method beyond this implementation detail.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Number of Attention Heads: 32 for Q and 2 for KV
coremodel architecturein Qwen3.5-122B-A10BQwen
Filed alongside
Other methods under model architecture :: token mixer :: linear attention & state space :: gated delta network.
Gated DeltaNetKimi Delta AttentionGated DeltaNet–sparse MoE hybridKDAGated DeltaNet with bounded sigmoid output gateHybrid linear attentionKimi Delta Attention and Attention Residuals architectureKimi Delta Attention with input-dependent full-rank output gateKimi Delta Attention with lower-bounded log-decaySimpleGDN