ambiguous · filed under model architecture
Hybrid Mamba-Attention
A hybrid architecture identified as combining Mamba with attention, without further mechanism details in the evidence.
- sources
- 3
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 4
Documented in
Evidence
4 spans quoted from the sources, strongest treatment first.
Mixture-of-Experts hybrid Mamba-Attention language model
coremodel architecturein Nemotron 3 UltraNVIDIA
Employs Mixture-of-Experts Hybrid Mamba-Attention architecture.
coremodel architecturein Nemotron 3 UltraNVIDIA
employing a Mixture-of-Experts hybrid Mamba-Attention architecture
coremodel architecturein Nemotron 3 UltraNVIDIA
Nemotron 3 Ultra uses a MoE Hybrid Mamba-Attention architecture
coremodel architecturein Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under model architecture :: token mixer :: hybrid layer stacking.
Hybrid AttentionMamba-2, Mixture-of-Experts, and Selective Attention HybridHybrid Mamba-TransformerGated DeltaNet and Gated AttentionHybrid Mamba-Transformer Mixture-of-Experts Layer LayoutDense Attention FallbackGated Attention and Sliding-Window Attention ConfigurationGated DeltaNet and Full AttentionGated DeltaNet and Qwen Sparse Attentionhybrid Gated DeltaNet + sparse MoE architectureHybrid sparse mixture-of-experts Transformer architectureSearch-Based Sliding-Window Attention Pattern