Model techniques map
Taxonomymodel architecturemultimodal architecture

taxonomy node · level 2

multimodal architecture

57 methods filed at this node or below it, from the sources of 16 models.

model architecture :: multimodal architecture

Matching aids for the classifier: vision encoder; audio encoder; lightweight projector; ViT; omnimodal architecture; language backbone.

In this branch 57

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 57

Encoder-free multimodal architecture core · 6 sources · 6 quotes
Early fusion multimodal training core · 5 sources · 5 quotes
Discrete token encoding for audio core · 4 sources · 4 quotes
Hierarchical patch encoder core · 4 sources · 4 quotes
Mixed-modality training core · 4 sources · 4 quotes
Native multimodal understanding core · 3 sources · 3 quotes
3×3 pixel-unshuffle downsampling core · 2 sources · 2 quotes
Joint vision-language pre-training core · 2 sources · 2 quotes
Joint multimodal input and understanding core · 1 source · 1 quote
Language backbone core · 1 source · 1 quote
Multimodal mixture-of-experts Transformer core · 1 source · 1 quote
Native multimodal encoding core · 1 source · 1 quote
Native multimodal visual understanding core · 1 source · 1 quote
Native visual and audio understanding core · 1 source · 1 quote
Single shared multimodal backbone core · 1 source · 1 quote
Thinker–Talker architecture core · 1 source · 1 quote
Vision encoder core · 1 source · 1 quote
Vision encoder and multimodal aligner core · 1 source · 1 quote
Vision encoder–MLP projector pathway core · 1 source · 1 quote
Vision Transformer (ViT) used · 2 sources · 2 quotes
2×2 pixel-shuffle downsampling used · 1 source · 1 quote
Audio encoder initialized from MiMo-Audio used · 1 source · 1 quote
Audio Transformer (AuT) used · 1 source · 1 quote
Causal streaming ConvNet codec decoder used · 1 source · 1 quote
Data-parallel-first multimodal encoding used · 1 source · 1 quote
Decoupled Encoder Process (DEP) used · 1 source · 1 quote
Dedicated multimodal encoders used · 1 source · 1 quote
Disaggregated encoder training used · 1 source · 1 quote
Dual-format coordinate supervision used · 1 source · 1 quote
Explicit text-string timestamps used · 1 source · 1 quote
Frozen encoders during pre-training used · 1 source · 1 quote
Joint full-model multimodal training used · 1 source · 1 quote
Lightweight vision embedding module used · 1 source · 1 quote
Linear-projection patch embedding used · 1 source · 1 quote
Mel-spectrogram audio front end used · 1 source · 1 quote
MoonViT-V2 used · 1 source · 1 quote
Multimodal input used · 1 source · 1 quote
Projector warmup used · 1 source · 1 quote
Reusing the image token for video frames used · 1 source · 1 quote
Staged vision-encoder freezing used · 1 source · 1 quote
Two-layer MLP vision projector used · 1 source · 1 quote
Variable aspect-ratio image handling used · 1 source · 1 quote
Configurable visual token budget optional · 3 sources · 3 quotes
Video preprocessor longest-edge configuration optional · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.

Modeltechniques
GLM-5.3-Flash usedReusing the image token for video frames used—
DeepSeek-V4.1-Flash core3×3 pixel-unshuffle downsampling coreJoint vision-language pre-training coreMultimodal mixture-of-experts Transformer coreVision encoder–MLP projector pathway coreDisaggregated encoder training usedDiscarding the LLM after vision-encoder training usedHigh-resolution autoregressive vision-encoder fine-tuning usedLinear-projection patch embedding usedStaged vision-encoder freezing usedTwo-layer MLP vision projector used—
DeepSeek-V4-Flash-0731 coreNative multimodal visual understanding core—
MiMo-V2.6-Flash coreFour-frame audio patches with within-patch bidirectional self-attention coreJoint multimodal input and understanding coreTwo-stage audio encoding with tokenization and patch encoding coreData-parallel-first multimodal encoding usedJoint full-model multimodal training usedTwo-stage text-then-multimodal pre-training used—
MiMo-V2.5 coreNative multimodal understanding coreNative visual and audio understanding coreAudio encoder initialized from MiMo-Audio usedProjector warmup usedVision Transformer (ViT) usedVisual and audio encoders with lightweight projectors used—
MiniMax-M3 coreMixed-modality training core—
DeepSeek-V4-Flash-Vision-Exp coreNative multimodal visual understanding coreVision encoder and multimodal aligner coreVisual modules for multimodal understanding core—
DeepSeek-V4-Pro-0813 coreNative multimodal visual understanding core—
Gemma 4 31B coreEncoder-free multimodal architecture coreDedicated multimodal encoders usedFrozen encoders during pre-training usedLightweight vision embedding module usedRaw audio projection into the LLM embedding space usedSingle-matrix-multiplication vision projection usedUniversal Speech Model (USM)-based audio encoder usedVariable aspect-ratio image handling usedConfigurable visual token budget optional—
Inkling coreDiscrete token encoding for audio coreHierarchical patch encoder coreJoint multimodal decoding in a shared hidden space coreNative multimodal encoding coreEncoder-free multimodal architecture used—
Kimi K3 coreSingle shared multimodal backbone core2×2 pixel-shuffle downsampling usedDecoupled Encoder Process (DEP) usedDual-format coordinate supervision usedFrom-scratch vision-encoder training with next-token prediction usedMoonViT-V2 usedMultimodal input used—
MiMo-V2.5-Pro coreNative multimodal understanding coreNative visual and audio understanding coreProjector warmup usedVisual and audio encoders with lightweight projectors used—
MiMo-V2.6-Pro coreFour-frame audio patches with within-patch bidirectional self-attention coreTwo-stage audio encoding with tokenization and patch encoding coreData-parallel-first multimodal encoding usedJoint full-model multimodal training usedTwo-stage text-then-multimodal pre-training used—
Qwen3.5-397B-A17B coreEarly fusion multimodal training coreThinker–Talker architecture coreAudio Transformer (AuT) usedCausal streaming ConvNet codec decoder usedExplicit text-string timestamps usedMel-spectrogram audio front end usedResidual Vector Quantization (RVQ) speech representation usedVideo preprocessor longest-edge configuration optional—
Qwen3.6-35B-A3B coreEarly fusion multimodal training core—
Step-3.7-Flash coreLanguage backbone coreVision encoder coreVision Transformer (ViT) used—