Model techniques map
Techniquesmodel architecturemultimodal architecture

specific method · filed under model architecture

Native multimodal understanding

Image, video, and audio understanding are combined natively in one model.

Also called native multimodal in one model, native omnimodal architecture, native omnimodal, native omnimodal model.

sources
3
models
2
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1core 2

Documented in

Evidence

3 spans quoted from the sources, strongest treatment first.

MiMo-V2.5 is natively omnimodal — trained from scratch to see, hear, and act across modalities

coremodel architecturein MiMo-V2.5Xiaomi

Native multimodal in one model. V2-Pro was text-and-code only. Multimodal capability existed in the separate V2-Omni model... V2.5 collapses both into a single architecture: native image, video, and audio understanding with Pro-level reasoning.

coremodel architecturein MiMo-V2.5-ProXiaomi

MiMo-V2.5 is a native omnimodal model by Xiaomi.

usedmodel architecturein MiMo-V2.5Xiaomi

Filed alongside

Other methods under model architecture :: multimodal architecture.