ambiguous · filed under model architecture
Main Branch
The MiniMax Sparse Attention branch that performs exact attention over selected blocks, though the evidence also describes it as using full attention.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
the Main Branch then performs exact block-sparse attention over only the selected blocks.
coremodel architecturein MiniMax-M3MiniMax
the Main Branch uses full attention
coremodel architecturein MiniMax Sparse AttentionMiniMax
Filed alongside
Other methods under model architecture :: token mixer :: sparse attention :: block selection.