implementation detail · filed under data curation
GitHub Crawl
A code-data crawl collected using the GitHub REST API and Amazon S3 API.
Also called GitHub Crawl via REST and S3 APIs.
- sources
- 2
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
The GitHub Crawl was collected using the GitHub REST API and the Amazon S3 API.
useddata curationin Nemotron 3.5 LightningNVIDIA
The GitHub Crawl was collected using the GitHub REST API and the Amazon S3 API
useddata curationin Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under data curation :: data sourcing.
Expert co-created training dataAgentic knowledge graph constructionCommon CrawlEssentialWebFinePDFsHigh-recall web-data curationImage-code pairs and computer-use trajectoriesLong-document curationNemotron-3-Ultra corpusNemotron-CCNemotron-Post-Training-v3Star-threshold-filtered GitHub repositoriesStreaming training-data ingestionVideo frame extraction at 1 FPSVideo sampling parametersWeb knowledge graph construction and question generation