Model techniques map
Techniquesmodel architecturetoken mixersoftmax attentionglobal attention

specific method · filed under model architecture

Multi-Head Attention (MHA)

The evidence identifies MHA as a reliability-related feature of a full-attention model, but does not establish a distinct attention mechanism beyond MHA.

Also called Multi-Head Attention (MHA) for reliability.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

MiniMax-M2.5 is the only full attention model (see this article) with Multi-Head Attention (MHA) for reliability purposes.

usedmodel architecturein MiniMax-M2.5MiniMax

Filed alongside

Other methods under model architecture :: token mixer :: softmax attention :: global attention.