Model techniques map
Techniquesmodel architecturetoken mixerhybrid layer stacking

specific method · filed under model architecture

Gated DeltaNet and Full Attention

Alternates Gated DeltaNet linear-attention layers with full-attention layers.

Also called Hybrid Attention Architecture.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

The model alternates between Gated DeltaNet layers (linear attention) and full attention layers in roughly a 3:1 ratio.

coremodel architecturein Qwen3.5-397B-A17BAlibaba

Filed alongside

Other methods under model architecture :: token mixer :: hybrid layer stacking.