Model techniques map
Techniquesdata curationdata mixture & curriculum

specific method · filed under data curation

Domain-specific multimodal datasets

Adds datasets targeting fine-grained visual perception, OCR, and long-tail knowledge.

Also called domain-specific datasets.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

we also incorporate domain-specific datasets to boost the model’s capabilities in fine-grained visual perception (e.g., visual grounding and pointing), optical character recognition (OCR), and the acquisition of long-tail knowledge.

usedunclearin DeepSeek-V4.1-FlashDeepSeek

Filed alongside

Other methods under data curation :: data mixture & curriculum.