specific method · filed under model architecture
Early fusion multimodal training
Multimodal tokens are included in training from the start, rather than added through a separate vision adapter afterward.
Also called Early fusion training on multimodal tokens, Unified Vision-Language Foundation, Early fusion vision-language training.
- sources
- 5
- models
- 2
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Further reading
Picked by hand, not extracted: where to read more, not evidence for anything on this page.
Evidence
5 spans quoted from the sources, strongest treatment first.
Early fusion training on multimodal tokens achieves cross-generational parity with Qwen3 and outperforms Qwen3-VL models across reasoning, coding, agents, and visual understanding benchmarks.
Early fusion training on multimodal tokens means the model doesn’t need a separate vision adapter.
Early fusion training on multimodal tokens achieves cross-generational parity with Qwen3 and outperforms Qwen3-VL models across reasoning, coding, agents, and visual understanding benchmarks.
Early fusion training on trillions of multimodal tokens achieves cross-generational parity with Qwen3 and outperforms Qwen3-VL models across reasoning, coding, agents, and visual understanding benchmarks.
Early fusion training on multimodal tokens achieves cross-generational parity with Qwen3
Filed alongside
Other methods under model architecture :: multimodal architecture.