specific method · filed under model architecture
Native multimodal understanding
Image, video, and audio understanding are combined natively in one model.
Also called native multimodal in one model, native omnimodal architecture, native omnimodal, native omnimodal model.
- sources
- 3
- models
- 2
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
MiMo-V2.5 is natively omnimodal — trained from scratch to see, hear, and act across modalities
Native multimodal in one model. V2-Pro was text-and-code only. Multimodal capability existed in the separate V2-Omni model... V2.5 collapses both into a single architecture: native image, video, and audio understanding with Pro-level reasoning.
MiMo-V2.5 is a native omnimodal model by Xiaomi.
Filed alongside
Other methods under model architecture :: multimodal architecture.