implementation detail · filed under evaluation
High-confidence ProgramBench subset
A ProgramBench subset retaining tasks whose reference solution achieves at least a 95% pass rate on the hidden test suite.
Also called High-confidence benchmark subset filtering, high-confidence subset of ProgramBench.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
We construct a high-confidence subset of ProgramBench by retaining only tasks for which the reference solution achieves a pass rate of at least 95% on the hidden test suite.
useddata curationin DeepSeek-V4.1-FlashDeepSeek
by retaining only tasks for which the reference solution achieves a pass rate of at least 95% on the hidden test suite
usedevaluation onlyin DeepSeek-V4.1-FlashDeepSeek
Filed alongside
Other methods under evaluation :: benchmark.
Artificial Analysis Intelligence IndexBrowseComp with SearchCAD visual reproductionCapture-the-Flag (CTF) evaluationComputer use closed-loop evaluationDisabling tool searchDomain whitelistExecution-grounded evaluationF2P/P2F acceptance criterionFail-to-pass and pass-to-pass evaluation pointsGPU kernel optimization task suiteKernelBenchLiving in-house benchmark suiteLongBench v2Multiple-rollout evaluationPaperBenchPinchBenchProfBench with SearchScaleAI Multi Challenge Multi Turn Instruction FollowingSQuADSWE-bench VerifiedThree-configuration cyber range testingVals.aiVerified CUDA kernels