specific method · filed under evaluation
Vulnerability discovery with proof-of-concept evaluation
Assessing whether a model identifies genuine bugs in current codebases and demonstrates that they are reproducible.
Also called Vulnerability discovery (Tier 1).
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
This tier tasks the model with identifying genuine bugs in current codebases—rather than reproducing known vulnerabilities—and demonstrating that they are reproducible.
usedevaluation onlyin Kimi K3 Tier 1 vulnerability discovery evaluationMoonshot AI
Filed alongside
Other methods under evaluation.
Prompt-based output standardizationEvaluation without safety filtersRefusal-suppressed variants for capability estimationavg@kAlmost@1Behavior testing with stubs and edge casesDeadline-bounded rollout evaluationDiverse benchmark validation for quantizationDocumenting evaluation configurationEnd-to-end exploit development evaluationEvaluation equivalence thresholdFixed MTP draft-token acceptance lengthFour-run mean pass@1Full trajectory releaseJoint validation protocolMean@5Model–harness co-designOfficial-release-source baseline restrictionPrompt engineering to elicit answersProtocolQA robustness validationRefusal behavior quantificationRollout-based auditingSample test and length-constraint filteringSandboxed GPU kernel optimization evaluation