specific method · filed under evaluation
Artificial Analysis Intelligence Index
An index combining nine evaluations to measure performance across agentic tasks, coding, scientific reasoning, and general intelligence.
- source
- 1
- model
- 1
- labs adopt it
- 0
- strongest
- evaluated
How sources treat it
One count per evidence span, weakest treatment to strongest.
evaluated 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
This index combines nine evaluations to measure model performance across agentic tasks, coding, scientific reasoning, and general intelligence.
evaluatedevaluation onlyin Nemotron 3.5 Lightning
Filed alongside
Other methods under evaluation :: benchmark.
High-confidence ProgramBench subsetBrowseComp with SearchCAD visual reproductionCapture-the-Flag (CTF) evaluationComputer use closed-loop evaluationDisabling tool searchDomain whitelistExecution-grounded evaluationF2P/P2F acceptance criterionFail-to-pass and pass-to-pass evaluation pointsGPU kernel optimization task suiteKernelBenchLiving in-house benchmark suiteLongBench v2Multiple-rollout evaluationPaperBenchPinchBenchProfBench with SearchScaleAI Multi Challenge Multi Turn Instruction FollowingSQuADSWE-bench VerifiedThree-configuration cyber range testingVals.aiVerified CUDA kernels