general family · filed under model architecture
Hybrid Mamba-Transformer
A Mamba-Transformer hybrid architecture, described in the evidence as a mixture-of-experts model in the named implementations.
Also called hybrid Mamba‑Transformer MoE, Mamba-Transformer.
- sources
- 4
- models
- 2
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Evidence
6 spans quoted from the sources, strongest treatment first.
Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
powered by hybrid Mamba‑Transformer MoE with 1M-token context
Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
the Nemotron 3 family of models use a Mixture-of-Experts hybrid Mamba-Transformer architecture that pushes the accuracy-to-inference-throughput frontier
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model
Filed alongside
Other methods under model architecture :: token mixer :: hybrid layer stacking.