Model techniques map
Taxonomyevaluationbenchmark

taxonomy node · level 2

benchmark

25 methods filed at this node or below it, from the sources of 9 models.

evaluation :: benchmark

Matching aids for the classifier: LongBench v2; SQuAD; HotpotQA; KernelBench; Tau Bench; BrowseComp; ProfBench; PinchBench.

In this branch 25

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 25

Living in-house benchmark suite core · 1 source · 1 quote
BrowseComp with Search used · 1 source · 1 quote
Capture-the-Flag (CTF) evaluation used · 1 source · 1 quote
Disabling tool search used · 1 source · 1 quote
Domain whitelist used · 1 source · 1 quote
Execution-grounded evaluation used · 1 source · 1 quote
F2P/P2F acceptance criterion used · 1 source · 1 quote
GPU kernel optimization task suite used · 1 source · 1 quote
High-confidence ProgramBench subset used · 1 source · 2 quotes
KernelBench used · 1 source · 1 quote
LongBench v2 used · 1 source · 1 quote
Multiple-rollout evaluation used · 1 source · 1 quote
PaperBench used · 1 source · 1 quote
PinchBench used · 1 source · 1 quote
ProfBench with Search used · 1 source · 1 quote
SQuAD used · 1 source · 1 quote
SWE-bench Verified used · 1 source · 1 quote
Three-configuration cyber range testing used · 1 source · 1 quote
Vals.ai used · 1 source · 1 quote
Verified CUDA kernels used · 1 source · 1 quote
CAD visual reproduction optional · 1 source · 1 quote
Computer use closed-loop evaluation optional · 1 source · 1 quote
Artificial Analysis Intelligence Index evaluated · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.