specific method · filed under evaluation
avg@k
Mean pass@1 averaged over multiple attempts per task, with the number of attempts specified by k.
Also called average over 3 rollouts (avg@3), avg@3.
- sources
- 2
- models
- 2
- labs adopt it
- 2
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
reporting avg@3 over three rollouts per task
usedunclearin GLM-5.3Z.ai
We report mean pass@1 averaged over multiple attempts per task (avg@k), with k per benchmark listed below.
usedevaluation onlyPoolside
Filed alongside
Other methods under evaluation.
Prompt-based output standardizationEvaluation without safety filtersRefusal-suppressed variants for capability estimationAlmost@1Behavior testing with stubs and edge casesDeadline-bounded rollout evaluationDiverse benchmark validation for quantizationDocumenting evaluation configurationEnd-to-end exploit development evaluationEvaluation equivalence thresholdFixed MTP draft-token acceptance lengthFour-run mean pass@1Full trajectory releaseJoint validation protocolMean@5Model–harness co-designOfficial-release-source baseline restrictionPrompt engineering to elicit answersProtocolQA robustness validationRefusal behavior quantificationRollout-based auditingSample test and length-constraint filteringSandboxed GPU kernel optimization evaluationSingle-agent and multi-agent evaluation configurations