specific method · filed under evaluation
Capture-the-Flag (CTF) evaluation
An evaluation in which a model solves CTF tasks using a headless Linux environment, offensive security tools, and a command-execution harness.
Also called CTF challenges.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
To evaluate the model against the CTFs we give it access to a headlessLinux distribution with common offensive cybersecurity tools preinstalled as well as a harness which allows the model to call those tools or otherwise execute commands similar to as a human.
usedevaluation onlyin gpt-oss-120b and gpt-oss-20bOpenAI
Filed alongside
Other methods under evaluation :: benchmark.
High-confidence ProgramBench subsetArtificial Analysis Intelligence IndexBrowseComp with SearchCAD visual reproductionComputer use closed-loop evaluationDisabling tool searchDomain whitelistExecution-grounded evaluationF2P/P2F acceptance criterionFail-to-pass and pass-to-pass evaluation pointsGPU kernel optimization task suiteKernelBenchLiving in-house benchmark suiteLongBench v2Multiple-rollout evaluationPaperBenchPinchBenchProfBench with SearchScaleAI Multi Challenge Multi Turn Instruction FollowingSQuADSWE-bench VerifiedThree-configuration cyber range testingVals.aiVerified CUDA kernels