Model techniques map
Techniquesmodel architecture

implementation detail · filed under model architecture

Short convolutions in attention and residual branches

Short convolutions applied after attention key and value projections and to attention and MLP residual-branch outputs.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

We also apply short convolutions at two points — after the key and value projections in each attention layer, and on the attention and MLP residual branch outputs before they rejoin the main residual stream.

usedmodel architecturein InklingThinking Machines Lab

Filed alongside

Other methods under model architecture.