taxonomy node · level 2
data sourcing
17 methods filed at this node or below it, from the sources of 10 models.
data curation :: data sourcing
Matching aids for the classifier: Common Crawl; GitHub crawl; web knowledge graph construction; Nemotron-CC; FinePDFs; EssentialWeb.
In this branch 17
Everything filed at this node or below it, with one collapsible heading per child node.
filed here 17
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| DeepSeek-V4.1-Flash used | Image-code pairs and computer-use trajectories usedStar-threshold-filtered GitHub repositories used— |
| Hy4-preview used | Expert co-created training data used— |
| NVIDIA-Nemotron-3-Ultra-550B-A55B used | Common Crawl usedEssentialWeb usedFinePDFs usedGitHub Crawl usedNemotron-3-Ultra corpus usedNemotron-CC usedNemotron-Post-Training-v3 used— |
| GLM-5.2 used | Web knowledge graph construction and question generation used— |
| DeepSeek-V4-Pro used | Long-document curation used— |
| Gemma 4 31B used | Video frame extraction at 1 FPS used— |
| Kimi K3 used | Agentic knowledge graph construction used— |
| Laguna-S-2.1 used | High-recall web-data curation usedStreaming training-data ingestion used— |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B used | GitHub Crawl used— |
| Qwen3.6-35B-A3B optional | Video sampling parameters optional— |