Model techniques map
Techniquesmodel architecturemultimodal architecture

general family · filed under model architecture

Native multimodal encoding

Images and video use a hierarchical patch encoder, audio uses discrete token encoding, and modalities are jointly processed by the decoder.

Also called natively multimodal.

source
1
model
1
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

The model is natively multimodal — images and video are encoded via a hierarchical patch encoder, and audio via discrete token encoding — with all modalities projected into a shared hidden space and processed jointly by the decoder.

coremodel architecturein InklingThinking Machines Lab

Filed alongside

Other methods under model architecture :: multimodal architecture.