specific method · filed under data curation
CBRN pre-training data filtering
Harmful pre-training content, especially hazardous biosecurity knowledge, is filtered using GPT-4o's CBRN filters.
Also called CBRN safety filtering of pre-training data, CBRN safety pre-training data filtering.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
we filtered the data for harmful content in pre-training, especially around hazardous biosecurity knowledge, by reusing the CBRN pre-training filters from GPT-4o
useddata curationin gpt-oss-120b and gpt-oss-20bOpenAI
we filtered the data for harmful content in pre-training, especially around hazardous biosecurity knowledge, by reusing the CBRN pre-training filters from GPT-4o.
useddata curationin gpt-oss-120b and gpt-oss-20bOpenAI
Filed alongside
Other methods under data curation :: data filtering.
CSAM filteringSensitive data filteringFiltering batched auto-generated and templated contentHeuristic filteringKeyword- and regex-based filteringModel-based refusal and inability filteringPass@100 filteringPathological repetition filteringSmolVLM-based image-text quality scoringUnified data filtering pipelineAutomated verificationCoding-agent session filtering and trajectory deduplicationConservative model-based noise filteringContent quality and safety filteringContinuous contribution-score rankingData cleaningData filtering pipelineDense annotation of ambiguous low-quality dataDomain filtering with heuristics, classifier scoring, and deduplicationEvidence-based data cleaning and training constraintsFiltering low-information model-generated contentHeuristic and statistical filtering with deduplication and quality modelsImage-text relevance filteringLong-context data cleaning pipeline