specific method · filed under model architecture
Mamba
A state-space sequence model, distinguished here from the separately named Mamba-2 variant.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
Nemotron 3 Ultra : Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
coremodel architecturein Nemotron 3 UltraNVIDIA
Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
coremodel architecturein Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under model architecture :: token mixer :: linear attention & state space :: Mamba.