Model techniques map
Taxonomydata curationdata sourcing

taxonomy node · level 2

data sourcing

17 methods filed at this node or below it, from the sources of 10 models.

data curation :: data sourcing

Matching aids for the classifier: Common Crawl; GitHub crawl; web knowledge graph construction; Nemotron-CC; FinePDFs; EssentialWeb.

In this branch 17

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 17

Expert co-created training data used · 2 sources · 2 quotes
GitHub Crawl used · 2 sources · 2 quotes
Agentic knowledge graph construction used · 1 source · 1 quote
Common Crawl used · 1 source · 1 quote
EssentialWeb used · 1 source · 1 quote
FinePDFs used · 1 source · 1 quote
High-recall web-data curation used · 1 source · 1 quote
Long-document curation used · 1 source · 1 quote
Nemotron-3-Ultra corpus used · 1 source · 1 quote
Nemotron-CC used · 1 source · 1 quote
Nemotron-Post-Training-v3 used · 1 source · 1 quote
Streaming training-data ingestion used · 1 source · 1 quote
Video frame extraction at 1 FPS used · 1 source · 1 quote
Video sampling parameters optional · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.