implementation detail · filed under model architecture
Short convolutions in attention and residual branches
Short convolutions applied after attention key and value projections and to attention and MLP residual-branch outputs.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
We also apply short convolutions at two points — after the key and value projections in each attention layer, and on the attention and MLP residual branch outputs before they rejoin the main residual stream.
usedmodel architecturein InklingThinking Machines Lab
Filed alongside
Other methods under model architecture.