Model techniques map
Techniquesdata curationdata filtering

specific method · filed under data curation

CBRN pre-training data filtering

Harmful pre-training content, especially hazardous biosecurity knowledge, is filtered using GPT-4o's CBRN filters.

Also called CBRN safety filtering of pre-training data, CBRN safety pre-training data filtering.

sources
2
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 2

Documented in

Evidence

2 spans quoted from the sources, strongest treatment first.

we filtered the data for harmful content in pre-training, especially around hazardous biosecurity knowledge, by reusing the CBRN pre-training filters from GPT-4o

useddata curationin gpt-oss-120b and gpt-oss-20bOpenAI

we filtered the data for harmful content in pre-training, especially around hazardous biosecurity knowledge, by reusing the CBRN pre-training filters from GPT-4o.

useddata curationin gpt-oss-120b and gpt-oss-20bOpenAI

Filed alongside

Other methods under data curation :: data filtering.