Model techniques map
Taxonomyevaluationhuman & real-world evaluation

taxonomy node · level 2

human & real-world evaluation

19 methods filed at this node or below it, from the sources of 13 models.

evaluation :: human & real-world evaluation

Matching aids for the classifier: internal R&D task evaluation; white-collar enterprise tasks; real-world feedback refinement.

In this branch 19

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 19

Held-out evaluation default · 1 source · 1 quote
Blind side-by-side expert evaluation used · 4 sources · 4 quotes
Multi-turn open-ended external red-teaming used · 3 sources · 3 quotes
Automated red-teaming loop used · 1 source · 1 quote
Cyber range exercises used · 1 source · 1 quote
Dangerous-capability uplift assessment used · 1 source · 1 quote
Human evaluation used · 1 source · 1 quote
Internal engineering task evaluation used · 1 source · 2 quotes
Pairwise comparison for Chinese writing used · 1 source · 1 quote
Pre-release safety evaluation used · 1 source · 1 quote
Real-world codebase testing used · 1 source · 1 quote
Real-world feedback refinement used · 1 source · 1 quote
Reward hacking mitigation in evaluations used · 1 source · 1 quote
Same-prompt comparative DevOps testing used · 1 source · 1 quote
White-collar enterprise task evaluation used · 1 source · 1 quote
A/B testing on real tasks mentioned · 1 source · 1 quote
Continuous monitoring mentioned · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.