Model techniques map
Techniquesmodel architecturemultimodal architecture

specific method · filed under model architecture

Universal Speech Model (USM)-based audio encoder

An audio encoder based on USM, with two downsampling convolution layers followed by twelve Conformer layers.

Also called Universal Speech Model-based audio encoder, USM-based audio encoder.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

The encoder architecture is based on the Universal Speech Model [Zhang et al., 2023, USM], consisting of two downsampling convolution layers followed by twelve Conformer layers [Gulati et al., 2020].

usedmodel architecturein Gemma 4Google DeepMind

Filed alongside

Other methods under model architecture :: multimodal architecture.