specific method · filed under model architecture
Mixed-modality training
Different modalities are mixed in training from the first step.
Also called mixed modalities from the start.
- sources
- 4
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 4
Documented in
Further reading
Picked by hand, not extracted: where to read more, not evidence for anything on this page.
- Chameleon: Mixed-Modal Early-Fusion Foundation Models paper arxiv.org
Evidence
4 spans quoted from the sources, strongest treatment first.
M3 is a model that has undergone mixed-modality training from Step 0.
coretraining objectivein MiniMax-M3MiniMax
M3 undergoes mixed-modality training from the very first step
coretraining objectivein MiniMax-M3MiniMax
MiniMax says M3 was trained with mixed modalities from the start.
coretraining objectivein MiniMax-M3MiniMax
M3 undergoes mixed-modality training from the very first step
coretraining objectivein MiniMax-M3MiniMax
Filed alongside
Other methods under model architecture :: multimodal architecture.
Encoder-free multimodal architectureEarly fusion multimodal trainingDiscrete token encoding for audioHierarchical patch encoderConfigurable visual token budgetJoint multimodal decoding in a shared hidden spaceNative multimodal understanding3×3 pixel-unshuffle downsamplingJoint vision-language pre-trainingRaw audio projection into the LLM embedding spaceVision Transformer (ViT)2×2 pixel-shuffle downsamplingAudio encoder initialized from MiMo-AudioAudio Transformer (AuT)Causal streaming ConvNet codec decoderData-parallel-first multimodal encodingDecoupled Encoder Process (DEP)Dedicated multimodal encodersDisaggregated encoder trainingDiscarding the LLM after vision-encoder trainingDual-format coordinate supervisionExplicit text-string timestampsFour-frame audio patches with within-patch bidirectional self-attentionFrom-scratch vision-encoder training with next-token prediction