Model techniques map
Techniquesmodel architecturetoken mixersoftmax attentionsliding window attention

specific method · filed under model architecture

512-token Sliding Window Attention

Sliding Window Attention using a 512-token window.

Also called + SWA-512.

sources
2
models
2
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

evaluated 1used 1

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

Of the 66 decoder layers, 55 use local attention with a small 512-token window.

usedmodel architecturein InklingThinking Machines Lab

+ SWA-512

evaluatedmodel architecturein Laguna XS.2Poolside

Filed alongside

Other methods under model architecture :: token mixer :: softmax attention :: sliding window attention.