Model techniques map
Taxonomymodel architecture

taxonomy area · pipeline stage 03

model architecture

312 methods filed at this node or below it, from the sources of 27 models.

model architecture

Matching aids for the classifier: what the network is, layer by layer.

In this branch 312

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 7

Causal Encoder-Decoder (CED) architecture core · 4 sources · 7 quotes
Per-Layer Embeddings (PLE) core · 4 sources · 4 quotes
Asymmetric architecture core · 1 source · 1 quote
Regular Transformer decoder core · 1 source · 1 quote
Dense model architecture used · 1 source · 1 quote

token mixer 123

Cross-layer attention sharing core · 1 source · 1 quote

softmax attention 32

Attention layers core · 1 source · 1 quote
DFlash attention core · 1 source · 1 quote
Softplus attention gating core · 1 source · 1 quote
Learned softmax denominator bias used · 1 source · 1 quote
Cascade attention not used · 2 sources · 2 quotes

global attention 4

Global attention core · 4 sources · 4 quotes
Unified Keys and Values core · 2 sources · 2 quotes
Full attention used · 2 sources · 2 quotes
Multi-Head Attention (MHA) used · 1 source · 1 quote

sliding window attention 12

Sliding Window Attention core · 8 sources · 10 quotes
Hybrid Sliding Window Attention core · 4 sources · 5 quotes
Sink-Augmented Sliding Window Attention core · 1 source · 1 quote
512-token Sliding Window Attention used · 2 sources · 2 quotes
Decoder SWA Bounded Replay used · 1 source · 2 quotes
Dynamic Attention Window Size Training used · 1 source · 1 quote
Pure Sliding Window Attention used · 1 source · 1 quote
SWA-128 used · 1 source · 1 quote
Fixed Sliding Window Attention evaluated · 1 source · 1 quote
Per-Head Gating evaluated · 1 source · 1 quote
SWA-1024 evaluated · 1 source · 1 quote

grouped-query attention 4

Grouped-query attention core · 11 sources · 13 quotes
Block-sparse grouped-query attention core · 1 source · 1 quote
Multi-query attention used · 2 sources · 2 quotes
Softplus-based per-head gating used · 1 source · 1 quote

multi-head latent attention 4

Gated Multi-Head Latent Attention core · 3 sources · 3 quotes
MQA Mode of Multi-Head Latent Attention core · 2 sources · 2 quotes
Multi-Head Latent Attention core · 2 sources · 3 quotes
Multi-Head Latent Attention MHA/MQA Modes used · 1 source · 1 quote

attention sink 3

Learnable attention sink bias used · 5 sources · 5 quotes
Attention sink used · 1 source · 1 quote
RoPE with attention sink unclear · 1 source · 1 quote

sparse attention 57

DeepSeek Sparse Attention core · 15 sources · 17 quotes
Compressed Sparse Attention core · 6 sources · 6 quotes
Qwen Sparse Attention core · 6 sources · 6 quotes
Sparse attention core · 6 sources · 6 quotes
Heavily Compressed Attention core · 5 sources · 5 quotes
Gated DeepSeek Sparse Attention core · 3 sources · 3 quotes
Fixed-budget sparse-attention selection core · 1 source · 1 quote
NoPE sparse multi-head latent attention core · 1 source · 1 quote
QSA micro-block compression at ratio 4 core · 1 source · 1 quote
Sequential block processing core · 1 source · 1 quote
Token-wise compression core · 1 source · 1 quote
KV-outer sparse attention used · 2 sources · 2 quotes
CSA2 Full Mode used · 1 source · 2 quotes
Natively trained sparsity used · 1 source · 1 quote
Sparse softmax attention used · 1 source · 1 quote
Two-stage introduction of sparse attention used · 1 source · 1 quote
Two-stage sparse attention used · 1 source · 1 quote
Sparse retrieval over long contexts evaluated · 1 source · 1 quote
Native Sparse Attention not used · 1 source · 1 quote

block selection 9

MiniMax Sparse Attention core · 6 sources · 6 quotes
Block-level selection core · 1 source · 1 quote
Dynamic sparse selection core · 1 source · 1 quote
Main Branch core · 1 source · 2 quotes
Top-512 block selection core · 1 source · 1 quote
Top-k block selection used · 2 sources · 2 quotes
Block Max Pooling used · 1 source · 1 quote
Top-scoring compressed-block selection used · 1 source · 1 quote
Mixture of Block Attention evaluated · 1 source · 1 quote

sparse attention indexer 22

IndexCache core · 3 sources · 3 quotes
IndexShare core · 3 sources · 3 quotes
Lightning Indexer core · 3 sources · 3 quotes
Compressed Sparse Attention 2 core · 2 sources · 6 quotes
Fine-grained token selection core · 2 sources · 2 quotes
Hierarchical Sparse Indexer core · 2 sources · 5 quotes
Compressed lightweight indexer core · 1 source · 1 quote
Index Branch core · 1 source · 2 quotes
IndexPool core · 1 source · 1 quote
MQA indexer core · 1 source · 1 quote
Reindex Mode core · 1 source · 3 quotes
Reuse Mode core · 1 source · 3 quotes
Dense Warm-up Stage used · 2 sources · 2 quotes
Average pooling used · 1 source · 1 quote
Block-causal scoring used · 1 source · 1 quote
Detached indexer-input optimization used · 1 source · 1 quote
Indexer Warmup used · 1 source · 2 quotes
Single-head index key used · 1 source · 1 quote
Index Branch output not used · 1 source · 1 quote
Index Branch value head not used · 1 source · 1 quote

fixed-pattern sparse attention 2

Sparse Global Attention Anchors core · 1 source · 1 quote
Local Block used · 1 source · 2 quotes

linear attention & state space 16

Linear attention core · 2 sources · 2 quotes
Lightning Attention not used · 1 source · 1 quote

Mamba 3

Mamba-2 core · 5 sources · 6 quotes
Mamba core · 2 sources · 2 quotes
Mamba-2 SSM cache optional · 1 source · 1 quote

gated delta network 11

Gated DeltaNet core · 13 sources · 13 quotes
Kimi Delta Attention core · 6 sources · 6 quotes
Gated DeltaNet–sparse MoE hybrid core · 2 sources · 2 quotes
KDA core · 2 sources · 2 quotes
Hybrid linear attention core · 1 source · 1 quote
SimpleGDN evaluated · 1 source · 1 quote

hybrid layer stacking 13

Hybrid Attention core · 27 sources · 29 quotes
Hybrid Mamba-Transformer core · 4 sources · 6 quotes
Hybrid Mamba-Attention core · 3 sources · 4 quotes
Gated DeltaNet and Full Attention core · 1 source · 1 quote
Gated DeltaNet and Gated Attention core · 1 source · 2 quotes
Gated DeltaNet and Qwen Sparse Attention core · 1 source · 1 quote
Search-Based Sliding-Window Attention Pattern evaluated · 1 source · 1 quote
Dense Attention Fallback unclear · 1 source · 1 quote

channel mixer 68

dense feed-forward network 6

SiTU-GLU core · 2 sources · 3 quotes
SwiGLU used · 3 sources · 3 quotes
SwiGLU clamping used · 3 sources · 3 quotes
Dense feed-forward network used · 1 source · 1 quote
SiTU used · 1 source · 1 quote
SwiGLU clipping not used · 1 source · 1 quote

mixture of experts 61

Mixture of Experts core · 88 sources · 101 quotes
Sparse expert activation core · 8 sources · 8 quotes
DeepSeekMoE core · 2 sources · 2 quotes
Mixture-of-Experts layers core · 2 sources · 2 quotes
MoE with routed and shared experts core · 2 sources · 2 quotes
Asymmetric input/output activation split core · 1 source · 1 quote
Dispatch recomputation core · 1 source · 1 quote
Gated DeltaNet MoE core · 1 source · 2 quotes
Hybrid Mixture of Experts core · 1 source · 1 quote
MegaMoE used · 1 source · 1 quote
MoE with 256 experts and top-8 routing used · 1 source · 1 quote
Routed expert output modulation used · 1 source · 1 quote

expert routing 15

Top-8 expert routing core · 5 sources · 5 quotes
Token-level expert routing core · 3 sources · 3 quotes
Keep Routing core · 1 source · 1 quote
Routed experts core · 1 source · 1 quote
Token-choice routing with softplus gating core · 1 source · 1 quote
Top-6 expert routing core · 1 source · 1 quote
Fixed Top-k routing with frozen bias default · 1 source · 1 quote
Anticipatory Routing used · 2 sources · 2 quotes
DP-aware routing used · 1 source · 1 quote
Hash routing used · 1 source · 1 quote
Latent-space routing used · 1 source · 1 quote
Token-choice routing used · 1 source · 1 quote
Top-4 expert routing used · 1 source · 2 quotes
Loss-spike-triggered Anticipatory Routing optional · 1 source · 1 quote

expert load balancing 15

Auxiliary-loss-free load balancing core · 3 sources · 5 quotes
Quantile Balancing core · 3 sources · 5 quotes
Auxiliary-loss load balancing used · 1 source · 1 quote
Exact coordinate minimization used · 1 source · 1 quote
Expert bias update factor used · 1 source · 1 quote
Expert Parallelism Load Balancing used · 1 source · 1 quote
formHC used · 1 source · 1 quote
Histogram-based quantile estimation used · 1 source · 2 quotes
Load balancing used · 1 source · 1 quote
Persistent load balancing used · 1 source · 1 quote
Pooled-global-batch quantile estimation used · 1 source · 1 quote
Round-robin load balancing optional · 1 source · 1 quote
MaxVio evaluated · 1 source · 1 quote

shared experts 7

Shared Experts core · 7 sources · 7 quotes
No Shared Experts core · 3 sources · 3 quotes
Shared Experts Active on Every Token core · 2 sources · 2 quotes
MoE with Expert Sinks core · 1 source · 1 quote

fine-grained experts 2

Granular Mixture of Experts not used · 2 sources · 2 quotes

latent mixture of experts 5

LatentMoE core · 7 sources · 9 quotes
Normalized LatentMoE core · 6 sources · 7 quotes
Hybrid Latent Mixture of Experts core · 1 source · 1 quote
Latent-dimension reduction used · 1 source · 1 quote

positional encoding 14

Relative attention core · 2 sources · 2 quotes
No Position Encoding core · 1 source · 2 quotes
Omitting RoPE in attention layers core · 1 source · 1 quote
Per-layer-type rotary position scales core · 1 source · 1 quote
YaRN used · 12 sources · 19 quotes
Rotary Position Embedding used · 5 sources · 5 quotes
2D rotary position embedding used · 3 sources · 3 quotes
Proportional Rotary Position Embedding used · 2 sources · 2 quotes
2D coordinate-based positional embeddings used · 1 source · 1 quote
Partial RoPE used · 1 source · 1 quote
RoPE scaling optional · 3 sources · 3 quotes
Gated attention with partial RoPE evaluated · 1 source · 1 quote
Temporal-Modality Rotary Position Embedding not used · 1 source · 1 quote

normalization & residual 28

Manifold-Constrained Hyper-Connections core · 8 sources · 8 quotes
Attention Residuals (AttnRes) core · 7 sources · 7 quotes
Gated Residual core · 6 sources · 8 quotes
Identity Hyper-Connections core · 3 sources · 3 quotes
Hyper-Connections core · 2 sources · 2 quotes
Per-Branch Scalar Write Gate core · 2 sources · 2 quotes
Single-Pass mHC core · 2 sources · 3 quotes
All-branch Elementwise Read Gate core · 1 source · 1 quote
Block AttnRes core · 1 source · 2 quotes
Four-Branch Residual Stream core · 1 source · 1 quote
GatedNorm core · 1 source · 3 quotes
Independent Per-Branch Normalization core · 1 source · 2 quotes
Low-Rank Gated Mix core · 1 source · 1 quote
Two-Phase Block AttnRes Schedule core · 1 source · 1 quote
Widened Residual Stream core · 1 source · 1 quote
RMSNorm used · 3 sources · 3 quotes
Bounded Positive Gates used · 1 source · 1 quote
Post-Embedding RMSNorm used · 1 source · 1 quote
Pre-LN used · 1 source · 1 quote
Zero-Centered RMSNorm used · 1 source · 1 quote
Simplified AltUp evaluated · 1 source · 1 quote
Full Attention Residuals mentioned · 1 source · 1 quote
QK-Clip not used · 2 sources · 2 quotes
Sparse Gated Residual Writes not used · 1 source · 1 quote

prediction head 5

Multi-Token Prediction core · 30 sources · 40 quotes
Nemotron Hybrid Multi-Token Prediction optional · 1 source · 1 quote

multimodal architecture 57

Encoder-free multimodal architecture core · 6 sources · 6 quotes
Early fusion multimodal training core · 5 sources · 5 quotes
Discrete token encoding for audio core · 4 sources · 4 quotes
Hierarchical patch encoder core · 4 sources · 4 quotes
Mixed-modality training core · 4 sources · 4 quotes
Native multimodal understanding core · 3 sources · 3 quotes
3×3 pixel-unshuffle downsampling core · 2 sources · 2 quotes
Joint vision-language pre-training core · 2 sources · 2 quotes
Joint multimodal input and understanding core · 1 source · 1 quote
Language backbone core · 1 source · 1 quote
Multimodal mixture-of-experts Transformer core · 1 source · 1 quote
Native multimodal encoding core · 1 source · 1 quote
Native multimodal visual understanding core · 1 source · 1 quote
Native visual and audio understanding core · 1 source · 1 quote
Single shared multimodal backbone core · 1 source · 1 quote
Thinker–Talker architecture core · 1 source · 1 quote
Vision encoder core · 1 source · 1 quote
Vision encoder and multimodal aligner core · 1 source · 1 quote
Vision encoder–MLP projector pathway core · 1 source · 1 quote
Vision Transformer (ViT) used · 2 sources · 2 quotes
2×2 pixel-shuffle downsampling used · 1 source · 1 quote
Audio encoder initialized from MiMo-Audio used · 1 source · 1 quote
Audio Transformer (AuT) used · 1 source · 1 quote
Causal streaming ConvNet codec decoder used · 1 source · 1 quote
Data-parallel-first multimodal encoding used · 1 source · 1 quote
Decoupled Encoder Process (DEP) used · 1 source · 1 quote
Dedicated multimodal encoders used · 1 source · 1 quote
Disaggregated encoder training used · 1 source · 1 quote
Dual-format coordinate supervision used · 1 source · 1 quote
Explicit text-string timestamps used · 1 source · 1 quote
Frozen encoders during pre-training used · 1 source · 1 quote
Joint full-model multimodal training used · 1 source · 1 quote
Lightweight vision embedding module used · 1 source · 1 quote
Linear-projection patch embedding used · 1 source · 1 quote
Mel-spectrogram audio front end used · 1 source · 1 quote
MoonViT-V2 used · 1 source · 1 quote
Multimodal input used · 1 source · 1 quote
Projector warmup used · 1 source · 1 quote
Reusing the image token for video frames used · 1 source · 1 quote
Staged vision-encoder freezing used · 1 source · 1 quote
Two-layer MLP vision projector used · 1 source · 1 quote
Variable aspect-ratio image handling used · 1 source · 1 quote
Configurable visual token budget optional · 3 sources · 3 quotes
Video preprocessor longest-edge configuration optional · 1 source · 1 quote

context capacity 10

N-gram embedding lookup core · 6 sources · 6 quotes
1M-token context window core · 4 sources · 4 quotes
Engram core · 3 sources · 5 quotes
256K context window core · 1 source · 1 quote
Single-layer N-gram embedding placement core · 1 source · 1 quote
Progressive context extension used · 2 sources · 2 quotes
Multi-head hashing used · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modelfiled heretoken mixerchannel mixerpositional encodingnormalization & residualprediction headmultimodal architecturecontext capacity
GLM-5.3-Flash core—Hybrid Attention coreIndexPool coreKimi Delta Attention coreNoPE sparse multi-head latent attention core—Mixture of Experts coreToken-level expert routing core——Manifold-Constrained Hyper-Connections core——Reusing the image token for video frames used——
DeepSeek-V4.1-Flash coreAsymmetric architecture coreCausal Encoder-Decoder (CED) architecture core—Compressed Sparse Attention 2 coreCross-layer attention sharing coreCross-layer KV and index reuse with statically assigned CSA2 modes coreHierarchical Sparse Indexer coreReindex Mode coreReuse Mode coreSliding Window Attention coreCross-stage shared-state management for attention reuse usedCSA2 Full Mode usedDecoder SWA Bounded Replay usedFrom-scratch sparse attention training without dense warmup usedMicro-batch-level shared-state lifetime management usedProgressive sequence-length extension for sparse attention usedPure Sliding Window Attention usedSparse retrieval over long contexts evaluated—Asymmetric input/output activation split coreAuxiliary-loss-free load balancing coreDeepSeekMoE shared and fine-grained routed experts coreMixture of Experts coreModality-specific auxiliary-loss-free load balancing coreMoE with 384 routed experts and a shared expert coreformHC usedMixture-of-Experts layers usedShared Experts usedSwiGLU usedSwiGLU clamping used—2D rotary position embedding used—Single-Pass mHC coreRMSNorm used——3×3 pixel-unshuffle downsampling coreJoint vision-language pre-training coreMultimodal mixture-of-experts Transformer coreVision encoder–MLP projector pathway coreDisaggregated encoder training usedDiscarding the LLM after vision-encoder training usedHigh-resolution autoregressive vision-encoder fine-tuning usedLinear-projection patch embedding usedStaged vision-encoder freezing usedTwo-layer MLP vision projector used—Engram coreRow-wise distributed partitioning of embedding tables used—
Hy4-preview core—Gated DeepSeek Sparse Attention coreIndexCache core—Mixture of Experts coreMoE with routed and shared experts coreShared Experts coreTop-8 expert routing coreTop-8 Routed-Expert Selection with Shared-Expert Activation core——Identity Hyper-Connections core—Multi-Token Prediction core———
DeepSeek-V4-Flash-0731 core——————Native multimodal visual understanding core——
NVIDIA-Nemotron-3-Ultra-550B-A55B core—Attention layers coreHybrid Mamba-Attention coreHybrid Mamba-Transformer coreMamba coreMamba-2 coreSparse Global Attention Anchors coreMamba-2 SSM cache optional—Hybrid Latent Mixture of Experts coreLatentMoE coreMixture of Experts coreMixture-of-Experts layers coreExpert Parallelism Load Balancing usedShared Experts usedMaxVio evaluatedGranular Mixture of Experts not used———Multi-Token Prediction coreSampled prior-MTP hidden-state conditioning usedShared-weight design across prediction heads usedNemotron Hybrid Multi-Token Prediction optional———
MiMo-V2.6-Flash core—Alternating Row-Major and Column-Major Token Serialization coreHybrid Attention coreHybrid sparse mixture-of-experts Transformer architecture coreSink-Augmented Sliding Window Attention core—Mixture of Experts coreNo Shared Experts core———Multi-Token Prediction core—Four-frame audio patches with within-patch bidirectional self-attention coreJoint multimodal input and understanding coreTwo-stage audio encoding with tokenization and patch encoding coreData-parallel-first multimodal encoding usedJoint full-model multimodal training usedTwo-stage text-then-multimodal pre-training used—1M-token context window core—
DeepSeek-V4-Flash core—Compressed Sparse Attention coreDeepSeek Sparse Attention coreHeavily Compressed Attention coreToken-wise compression coreAttention sink usedHybrid Attention usedMulti-query attention usedSliding Window Attention usedTwo-stage introduction of sparse attention used—DeepSeekMoE coreMixture of Experts coreShared Experts coreAuxiliary-loss-free load balancing usedHash routing usedMegaMoE usedSwiGLU usedSwiGLU clamping usedAnticipatory Routing optionalLoss-spike-triggered Anticipatory Routing optional—Rotary Position Embedding used—Manifold-Constrained Hyper-Connections coreRMSNorm usedQK-Clip not used—Multi-Token Prediction core——1M-token context window default—
MiMo-V2.5 core—Global attention coreHybrid Attention coreHybrid Sliding Window Attention coreSliding Window Attention coreFull attention usedGrouped-query attention usedLearnable attention sink bias usedSWA-128 used—Mixture of Experts coreDense feed-forward network usedExpert bias update factor usedLoad balancing usedMoE with 256 experts and top-8 routing usedTop-8 expert routing usedRound-robin load balancing optionalNo Shared Experts not used—Rotary Position Embedding used——Multi-Token Prediction core—Native multimodal understanding coreNative visual and audio understanding coreAudio encoder initialized from MiMo-Audio usedProjector warmup usedVision Transformer (ViT) usedVisual and audio encoders with lightweight projectors used—1M-token context window usedProgressive context extension used—
Hy3 core—Grouped-query attention core—Mixture of Experts coreRouted experts coreShared Experts coreTop-8 expert routing core———Multi-Token Prediction core———
GLM-5.2 core—DeepSeek Sparse Attention coreIndexShare coreMulti-Head Latent Attention coreLightning Indexer usedSparse attention usedGated DeltaNet evaluatedSearch-Based Sliding-Window Attention Pattern evaluatedSimpleGDN evaluatedSliding Window Attention evaluatedMulti-query attention mentioned—Mixture of Experts coreDP-aware routing used———Multi-Token Prediction used———
MiniMax-M3 core—Block-level selection coreBlock-sparse grouped-query attention coreDynamic sparse selection coreGrouped-query attention coreIndex Branch coreMain Branch coreMiniMax Sparse Attention coreSequential block processing coreSparse attention coreBlock Max Pooling usedIndexer Warmup usedKV-outer sparse attention usedLocal Block usedNatively trained sparsity usedSingle-head index key usedSparse softmax attention usedTop-k block selection usedTwo-stage sparse attention usedDeepSeek Sparse Attention evaluatedFixed Sliding Window Attention evaluatedFull attention evaluatedMixture of Block Attention evaluatedLinear attention mentionedSliding Window Attention mentionedCompressed Sparse Attention not usedDense Attention Fallback unclearIndex Branch output not usedIndex Branch value head not usedLearnable attention sink bias not usedLightning Attention not usedMulti-Head Latent Attention not usedNative Sparse Attention not usedRoPE with attention sink unclear—Mixture of Experts corePersistent load balancing usedShared Experts usedTop-4 expert routing used—Rotary Position Embedding used———Mixed-modality training core——
DeepSeek-V3.2 core—DeepSeek Sparse Attention coreFine-grained token selection coreLightning Indexer coreMQA Mode of Multi-Head Latent Attention coreDense Warm-up Stage usedDetached indexer-input optimization usedMulti-Head Latent Attention MHA/MQA Modes usedSparse-attention continued pre-training with joint model and indexer optimization used—Keep Routing core——————
DeepSeek-V4-Flash-Vision-Exp core—DFlash attention core—Mixture of Experts core——Hyper-Connections core——Native multimodal visual understanding coreVision encoder and multimodal aligner coreVisual modules for multimodal understanding core——
DeepSeek-V4-Pro core—Compressed Sparse Attention coreDeepSeek Sparse Attention coreHeavily Compressed Attention coreToken-wise compression coreAttention sink usedHybrid Attention usedMulti-query attention usedSliding Window Attention usedTwo-stage introduction of sparse attention used—DeepSeekMoE coreMixture of Experts coreShared Experts coreAnticipatory Routing usedAuxiliary-loss-free load balancing usedHash routing usedMegaMoE usedSwiGLU usedSwiGLU clamping usedLoss-spike-triggered Anticipatory Routing optional—Rotary Position Embedding used—Manifold-Constrained Hyper-Connections coreRMSNorm usedQK-Clip not used—Multi-Token Prediction core——1M-token context window defaultEngram not used—
DeepSeek-V4-Pro-0813 core——————Native multimodal visual understanding core——
Gemma 4 31B coreEncoder-free decoder-only Transformer architecture corePer-Layer Embeddings (PLE) coreDense model architecture used—Hybrid Attention coreUnified Keys and Values core—Mixture of Experts coreSparse expert activation core—2D coordinate-based positional embeddings used2D rotary position embedding usedProportional Rotary Position Embedding used———Encoder-free multimodal architecture coreDedicated multimodal encoders usedFrozen encoders during pre-training usedLightweight vision embedding module usedRaw audio projection into the LLM embedding space usedSingle-matrix-multiplication vision projection usedUniversal Speech Model (USM)-based audio encoder usedVariable aspect-ratio image handling usedConfigurable visual token budget optional——
Inkling coreRegular Transformer decoder coreShort convolutions in attention and residual branches used—Hybrid Attention core512-token Sliding Window Attention usedGrouped-query attention usedKernel-4 convolutional layers in decoder blocks used—Mixture of Experts coreMoE with Expert Sinks coreShared Experts Active on Every Token coreToken-level expert routing coreTop-6 expert routing coreSigmoid-based MoE router with auxiliary-loss-free load balancing used—Relative attention coreLearned input-dependent relative position bias used—Post-Embedding RMSNorm used——Discrete token encoding for audio coreHierarchical patch encoder coreJoint multimodal decoding in a shared hidden space coreNative multimodal encoding coreEncoder-free multimodal architecture used——
Kimi K3 core—Gated Multi-Head Latent Attention coreHybrid Attention coreHybrid linear attention coreKDA coreKimi Delta Attention coreKimi Delta Attention and Attention Residuals architecture coreFactorized spatial-temporal video attention with temporal pooling usedGated DeltaNet usedKimi Delta Attention with input-dependent full-rank output gate usedKimi Delta Attention with lower-bounded log-decay used—Auxiliary-loss-free load balancing coreDispatch recomputation coreLatentMoE coreMixture of Experts coreNormalized LatentMoE coreOutput-independent MoE gradient reformulation coreQuantile Balancing coreSharded latent weights with fused all-gather GEMM epilogue coreSiTU-GLU coreSparse expert activation coreFixed Top-k routing with frozen bias defaultExact coordinate minimization usedExponential moving average of estimated quantiles usedHistogram-based quantile estimation usedLatent-space routing usedPooled-global-batch quantile estimation usedSiTU used—No Position Encoding core—Attention Residuals (AttnRes) coreBlock AttnRes coreTwo-Phase Block AttnRes Schedule coreFull Attention Residuals mentioned——Single shared multimodal backbone core2×2 pixel-shuffle downsampling usedDecoupled Encoder Process (DEP) usedDual-format coordinate supervision usedFrom-scratch vision-encoder training with next-token prediction usedMoonViT-V2 usedMultimodal input used—Progressive context extension used—
Laguna-S-2.1 core—Grouped-query attention coreHybrid Attention coreSoftplus attention gating coreGated Attention and Sliding-Window Attention Configuration usedSoftplus-based per-head gating used512-token Sliding Window Attention evaluatedDense gated attention with full RoPE and full gating evaluatedPer-Head Gating evaluatedSWA-1024 evaluated—Mixture of Experts coreSparse expert activation coreToken-choice routing with softplus gating coreAuxiliary-loss load balancing usedRouted expert output modulation usedShared Experts usedToken-choice routing used—Per-layer-type rotary position scales coreRotary Position Embedding usedGated attention with partial RoPE evaluated—————
MiMo-V2.5-Pro core—Hybrid Attention coreGlobal attention usedHybrid Sliding Window Attention usedLearnable attention sink bias usedSliding Window Attention used—Mixture of Experts core———Multi-Token Prediction used—Native multimodal understanding coreNative visual and audio understanding coreProjector warmup usedVisual and audio encoders with lightweight projectors used—1M-token context window usedProgressive context extension used—
MiMo-V2.6-Pro core—Alternating Row-Major and Column-Major Token Serialization coreHybrid Attention coreHybrid sparse mixture-of-experts Transformer architecture coreSink-Augmented Sliding Window Attention core—Mixture of Experts coreNo Shared Experts coreSparse expert activation core———Multi-Token Prediction core—Four-frame audio patches with within-patch bidirectional self-attention coreTwo-stage audio encoding with tokenization and patch encoding coreData-parallel-first multimodal encoding usedJoint full-model multimodal training usedTwo-stage text-then-multimodal pre-training used——
NVIDIA-Nemotron-3.5-Lightning-30B-A3B core—Hybrid Mamba-Transformer coreMamba-2 coreMamba-2, Mixture-of-Experts, and Selective Attention Hybrid core—LatentMoE coreMixture of Experts coreToken-level expert routing coreLatent-dimension reduction usedMoE with 128 routed experts and a shared expert used—Omitting RoPE in attention layers core——Multi-Token Prediction core———
Qwen3.5-397B-A17B core—DeepSeek Sparse Attention coreGated DeltaNet coreGated DeltaNet and Full Attention coreGated DeltaNet with reduced KV-head configuration coreGated DeltaNet–sparse MoE hybrid corehybrid Gated DeltaNet + sparse MoE architecture coreHybrid Mamba-Transformer Mixture-of-Experts Layer Layout coreLinear attention coreDynamic Attention Window Size Training usedMulti-Head Attention (MHA) used—10 Routed + 1 Shared Experts Activated per Token core8 Routed + 1 Shared Experts Activated per Token coreHybrid Mixture of Experts coreMixture of Experts coreSparse expert activation core—YaRN usedRoPE scaling optionalTemporal-Modality Rotary Position Embedding not used——Multi-Token Prediction usedMulti-Token Prediction for residual codebooks used—Early fusion multimodal training coreThinker–Talker architecture coreAudio Transformer (AuT) usedCausal streaming ConvNet codec decoder usedExplicit text-string timestamps usedMel-spectrogram audio front end usedResidual Vector Quantization (RVQ) speech representation usedVideo preprocessor longest-edge configuration optional——
Qwen3.6-35B-A3B core—Gated DeltaNet coreGated DeltaNet and Gated Attention core—Gated DeltaNet MoE coreMixture of Experts core—RoPE scaling optionalYaRN optional———Early fusion multimodal training core——
Qwen3.8-Flash-Next core—Compressed lightweight indexer coreFixed-budget sparse-attention selection coreGated DeltaNet coreGated DeltaNet and Qwen Sparse Attention coreHybrid Attention coreMQA indexer coreQSA micro-block compression at ratio 4 coreQwen Sparse Attention coreReuse QSA index selection across speculative decoding steps coreTop-512 block selection coreAverage pooling usedBlock-causal scoring usedGated DeltaNet with bounded sigmoid output gate usedJoint backbone and indexer training under sparse attention usedTop-scoring compressed-block selection used—10 Routed + 1 Shared Experts Activated per Token coreMixture of Experts coreContextual gating for N-gram embedding injection usedFrequency-based partitioning of N-gram embedding slots evaluatedSwiGLU clipping not used—Partial RoPE usedYaRN used—All-branch Elementwise Read Gate coreDynamic Gating of Residual Reads and Writes coreElementwise Data-Dependent Residual Read Gate coreFour-Branch Residual Stream coreFour-Stream Hyper-Connection Combine Update coreGated Residual coreGatedNorm coreIndependent Per-Branch Normalization coreLow-Rank Gated Mix corePer-Branch Scalar Write Gate coreWidened Residual Stream coreBounded Positive Gates usedData-Dependent Residual Read and Write Operators usedZero-Centered RMSNorm usedHyper-Connections evaluatedManifold-Constrained Hyper-Connections evaluatedSimplified AltUp evaluatedQK-Clip not usedSparse Gated Residual Writes not used—Multi-Token Prediction core——N-gram embedding host-memory offload and prefetch coreN-gram embedding lookup coreSingle-layer N-gram embedding placement coreMulti-head hashing usedToken normalization for N-gram vocabulary compression evaluated—
Step-3.7-Flash core—Cascade attention not used—Mixture of Experts core———Multi-Token Prediction optional—Language backbone coreVision encoder coreVision Transformer (ViT) used—256K context window core—
gpt-oss-120b core—Grouped-query attention usedHybrid Attention usedLearned softmax denominator bias used—Mixture of Experts coreTop-k expert routing with softmax over selected experts coreSwiGLU used—Rotary Position Embedding usedYaRN used—Pre-LN usedRMSNorm used————