implementation detail · filed under model architecture
3×3 pixel-unshuffle downsampling
A visual-token downsampling operation that applies pixel-unshuffle at a 3×3 factor.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1core 1
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
we apply a pixel-unshuffle operation with 3× 3 downsampling to reduce the visual token count by a factor of nine
coremodel architecturein DeepSeek-V4.1-FlashDeepSeek
A vision encoder (DeepSeek-ViT, trained from scratch with 2D-RoPE and 3×3 pixel-unshuffle downsampling)
usedmodel architecturein DeepSeek-ViTDeepSeek
Filed alongside
Other methods under model architecture :: multimodal architecture.
Encoder-free multimodal architectureEarly fusion multimodal trainingDiscrete token encoding for audioHierarchical patch encoderMixed-modality trainingConfigurable visual token budgetJoint multimodal decoding in a shared hidden spaceNative multimodal understandingJoint vision-language pre-trainingRaw audio projection into the LLM embedding spaceVision Transformer (ViT)2×2 pixel-shuffle downsamplingAudio encoder initialized from MiMo-AudioAudio Transformer (AuT)Causal streaming ConvNet codec decoderData-parallel-first multimodal encodingDecoupled Encoder Process (DEP)Dedicated multimodal encodersDisaggregated encoder trainingDiscarding the LLM after vision-encoder trainingDual-format coordinate supervisionExplicit text-string timestampsFour-frame audio patches with within-patch bidirectional self-attentionFrom-scratch vision-encoder training with next-token prediction