specific method · filed under model architecture
Multi-Head Attention (MHA)
The evidence identifies MHA as a reliability-related feature of a full-attention model, but does not establish a distinct attention mechanism beyond MHA.
Also called Multi-Head Attention (MHA) for reliability.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
MiniMax-M2.5 is the only full attention model (see this article) with Multi-Head Attention (MHA) for reliability purposes.
usedmodel architecturein MiniMax-M2.5MiniMax
Filed alongside
Other methods under model architecture :: token mixer :: softmax attention :: global attention.