Model techniques map
Techniquesmodel architecturemultimodal architecture

general family · filed under model architecture

Two-stage audio encoding with tokenization and patch encoding

Audio is first tokenized and then encoded in patches.

source
1
models
2
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

Audio encoding in MiMo-V2.6 consists of two stages: audio tokenization with the Audio Tokenizer, followed by patch encoding.

coremodel architecturein MiMo-V2.6Xiaomi

Filed alongside

Other methods under model architecture :: multimodal architecture.