general family · filed under model architecture
Vision Transformer (ViT)
A transformer architecture used as a vision encoder.
Also called Vision Transformer, ViT vision encoder.
- sources
- 2
- models
- 2
- labs adopt it
- 2
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
Equipped with a 729M-param Vision Transformer (ViT) featuring hybrid window attention
usedmodel architecturein MiMo-V2.5Xiaomi
paired with a 1.88-billion-parameter ViT vision encoder
usedmodel architecturein Step 3.7 FlashStepFun
Filed alongside
Other methods under model architecture :: multimodal architecture.
Encoder-free multimodal architectureEarly fusion multimodal trainingDiscrete token encoding for audioHierarchical patch encoderMixed-modality trainingConfigurable visual token budgetJoint multimodal decoding in a shared hidden spaceNative multimodal understanding3×3 pixel-unshuffle downsamplingJoint vision-language pre-trainingRaw audio projection into the LLM embedding space2×2 pixel-shuffle downsamplingAudio encoder initialized from MiMo-AudioAudio Transformer (AuT)Causal streaming ConvNet codec decoderData-parallel-first multimodal encodingDecoupled Encoder Process (DEP)Dedicated multimodal encodersDisaggregated encoder trainingDiscarding the LLM after vision-encoder trainingDual-format coordinate supervisionExplicit text-string timestampsFour-frame audio patches with within-patch bidirectional self-attentionFrom-scratch vision-encoder training with next-token prediction