specific method · filed under evaluation
Prompt-based output standardization
Using prompts to standardize model outputs when benchmarking.
Also called Standardize Output Format, standardize model outputs, Standardize output format for benchmarking.
- sources
- 5
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 2used 3
Documented in
Evidence
5 spans quoted from the sources, strongest treatment first.
We recommend using prompts to standardize model outputs when benchmarking.
usedevaluation onlyin Qwen3.5-122B-A10BQwen
We recommend using prompts to standardize model outputs when benchmarking.
usedevaluation onlyQwen
We recommend using prompts to standardize model outputs when benchmarking.
usedevaluation onlyin Qwen3.6-27BQwen
We recommend using prompts to standardize model outputs when benchmarking.
optionalevaluation onlyin Qwen3.5-35B-A3BQwen
We recommend using prompts to standardize model outputs when benchmarking.
optionalevaluation onlyin Qwen3.6-35B-A3BQwen
Filed alongside
Other methods under evaluation.
Evaluation without safety filtersRefusal-suppressed variants for capability estimationavg@kAlmost@1Behavior testing with stubs and edge casesDeadline-bounded rollout evaluationDiverse benchmark validation for quantizationDocumenting evaluation configurationEnd-to-end exploit development evaluationEvaluation equivalence thresholdFixed MTP draft-token acceptance lengthFour-run mean pass@1Full trajectory releaseJoint validation protocolMean@5Model–harness co-designOfficial-release-source baseline restrictionPrompt engineering to elicit answersProtocolQA robustness validationRefusal behavior quantificationRollout-based auditingSample test and length-constraint filteringSandboxed GPU kernel optimization evaluationSingle-agent and multi-agent evaluation configurations