Model techniques map
Techniquesmodel architecturemultimodal architecture

general family · filed under model architecture

Single shared multimodal backbone

Text, images, and video are processed by one shared backbone within a context, without a post-hoc modality-alignment stage.

Also called single shared backbone.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

text, images, and videos are processed by a single shared backbone within one context, with no post-hoc modality-alignment stage.

coremodel architecturein Kimi K3Moonshot AI

Filed alongside

Other methods under model architecture :: multimodal architecture.