specific method · filed under evaluation
Computer use closed-loop evaluation
An evaluation in which a model inspects an open application and uses its observed structure and feedback to recreate a functional version.
Also called Computer use closed loop.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- optional
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Use Computer Use to inspect the currently open application, identify its page structure, core functionality, primary user flows, and interaction feedback, and recreate a functional version in the current directory.
optionalotherZ.ai
Filed alongside
Other methods under evaluation :: benchmark.
High-confidence ProgramBench subsetArtificial Analysis Intelligence IndexBrowseComp with SearchCAD visual reproductionCapture-the-Flag (CTF) evaluationDisabling tool searchDomain whitelistExecution-grounded evaluationF2P/P2F acceptance criterionFail-to-pass and pass-to-pass evaluation pointsGPU kernel optimization task suiteKernelBenchLiving in-house benchmark suiteLongBench v2Multiple-rollout evaluationPaperBenchPinchBenchProfBench with SearchScaleAI Multi Challenge Multi Turn Instruction FollowingSQuADSWE-bench VerifiedThree-configuration cyber range testingVals.aiVerified CUDA kernels