Model techniques map
Techniquesmodel architecturemultimodal architecture

implementation detail · filed under model architecture

Four-frame audio patches with within-patch bidirectional self-attention

Consecutive groups of four frames form audio patches processed with bidirectional self-attention confined to each patch.

source
1
models
2
lab adopt it
1
strongest
core

How sources treat it

One count per evidence span, weakest treatment to strongest.

core 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

Every four consecutive frames are grouped into an audio patch and processed by a Transformer with bidirectional self-attention confined to that patch.

coremodel architecturein MiMo-V2.6 Audio Patch EncoderXiaomi

Filed alongside

Other methods under model architecture :: multimodal architecture.