specific method · filed under evaluation
A/B testing on real tasks
Compares alternatives on three real tasks spanning agentic, long-context, and knowledge work.
Also called A/B testing on three real tasks, A/B on three real tasks.
- source
- 1
- model
- 1
- labs adopt it
- 0
- strongest
- mentioned
How sources treat it
One count per evidence span, weakest treatment to strongest.
mentioned 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
A/B on three real tasks — one agentic, one long-context, one knowledge.
mentionedevaluation only
Filed alongside
Other methods under evaluation :: human & real-world evaluation.
Blind side-by-side expert evaluationMulti-turn open-ended external red-teamingInternal engineering task evaluationAutomated red-teaming loopContinuous monitoringCyber range exercisesDangerous-capability uplift assessmentExternal expert review of safety evaluation methodologyHeld-out evaluationHuman evaluationPairwise comparison for Chinese writingPre-release safety evaluationReal-world codebase testingReal-world feedback refinementReward hacking mitigation in evaluationsSame-prompt comparative DevOps testingSystematic robustness-boundary characterizationWhite-collar enterprise task evaluation