Model techniques map
Techniquesmodel architecturetoken mixerlinear attention & state spacegated delta network

implementation detail · filed under model architecture

Gated DeltaNet with reduced KV-head configuration

A Gated Attention configuration with 32 query heads and 2 key/value heads; the evidence does not establish a distinct named method beyond this implementation detail.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

Number of Attention Heads: 32 for Q and 2 for KV

coremodel architecturein Qwen3.5-122B-A10BQwen

Filed alongside

Other methods under model architecture :: token mixer :: linear attention & state space :: gated delta network.