Model techniques map
Techniquesmodel architecturetoken mixerhybrid layer stacking

general family · filed under model architecture

Hybrid Mamba-Transformer

A Mamba-Transformer hybrid architecture, described in the evidence as a mixture-of-experts model in the named implementations.

Also called hybrid Mamba‑Transformer MoE, Mamba-Transformer.

sources
4
models
2
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 2core 4

Documented in

Evidence

6 spans quoted from the sources, strongest treatment first.

Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

coremodel architecturein Nemotron 3 UltraNVIDIA

powered by hybrid Mamba‑Transformer MoE with 1M-token context

coremodel architecturein Nemotron 3 familyNVIDIA

Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

coremodel architecturein Nemotron 3 UltraNVIDIA

the Nemotron 3 family of models use a Mixture-of-Experts hybrid Mamba-Transformer architecture that pushes the accuracy-to-inference-throughput frontier

coremodel architecturein Nemotron 3 familyNVIDIA

Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model

usedmodel architecturein Nemotron 3 UltraNVIDIA

Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model

usedmodel architecturein Nemotron 3 UltraNVIDIA

Filed alongside

Other methods under model architecture :: token mixer :: hybrid layer stacking.