taxonomy node · level 2
benchmark
25 methods filed at this node or below it, from the sources of 9 models.
evaluation :: benchmark
Matching aids for the classifier: LongBench v2; SQuAD; HotpotQA; KernelBench; Tau Bench; BrowseComp; ProfBench; PinchBench.
In this branch 25
Everything filed at this node or below it, with one collapsible heading per child node.
filed here 25
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| GLM-5.3-Flash optional | CAD visual reproduction optionalComputer use closed-loop evaluation optional— |
| DeepSeek-V4.1-Flash used | Fail-to-pass and pass-to-pass evaluation points usedHigh-confidence ProgramBench subset usedMultiple-rollout evaluation used— |
| NVIDIA-Nemotron-3-Ultra-550B-A55B used | BrowseComp with Search usedKernelBench usedLongBench v2 usedPinchBench usedProfBench with Search usedScaleAI Multi Challenge Multi Turn Instruction Following usedSQuAD usedVals.ai used— |
| GLM-5.3 used | Disabling tool search usedDomain whitelist used— |
| DeepSeek-V3.2 used | F2P/P2F acceptance criterion used— |
| Kimi K3 core | Living in-house benchmark suite coreGPU kernel optimization task suite used— |
| Laguna-S-2.1 used | Execution-grounded evaluation used— |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B used | Verified CUDA kernels usedArtificial Analysis Intelligence Index evaluated— |
| gpt-oss-120b used | Capture-the-Flag (CTF) evaluation usedPaperBench usedSWE-bench Verified usedThree-configuration cyber range testing used— |