implementation detail · filed under evaluation
Domain whitelist
An evaluation-environment restriction that permits access only to approved domains to prevent agent cheating.
Also called domain whitelist for agent environments.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
We also remove all Git-related information and apply a domain whitelist (allowing only essential domains such as pypi.org and deb.debian.org for basic tool installation) to prevent the agent from cheating.
usedunclearin GLM-5.3Z.ai
Filed alongside
Other methods under evaluation :: benchmark.
High-confidence ProgramBench subsetArtificial Analysis Intelligence IndexBrowseComp with SearchCAD visual reproductionCapture-the-Flag (CTF) evaluationComputer use closed-loop evaluationDisabling tool searchExecution-grounded evaluationF2P/P2F acceptance criterionFail-to-pass and pass-to-pass evaluation pointsGPU kernel optimization task suiteKernelBenchLiving in-house benchmark suiteLongBench v2Multiple-rollout evaluationPaperBenchPinchBenchProfBench with SearchScaleAI Multi Challenge Multi Turn Instruction FollowingSQuADSWE-bench VerifiedThree-configuration cyber range testingVals.aiVerified CUDA kernels