specific method · filed under evaluation
Held-out evaluation
Reserves benchmarks from development decisions and evaluates them only for final validation.
Also called held-out generalization gates.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- default
How sources treat it
One count per evidence span, weakest treatment to strongest.
default 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
These benchmarks were reserved for final validation: they were not used for training-time monitoring, checkpoint selection, or any other development decisions, and were evaluated only once after the final model was produced.
defaultevaluation onlyin Nemotron 3 UltraNVIDIA
Filed alongside
Other methods under evaluation :: human & real-world evaluation.
Blind side-by-side expert evaluationMulti-turn open-ended external red-teamingInternal engineering task evaluationA/B testing on real tasksAutomated red-teaming loopContinuous monitoringCyber range exercisesDangerous-capability uplift assessmentExternal expert review of safety evaluation methodologyHuman evaluationPairwise comparison for Chinese writingPre-release safety evaluationReal-world codebase testingReal-world feedback refinementReward hacking mitigation in evaluationsSame-prompt comparative DevOps testingSystematic robustness-boundary characterizationWhite-collar enterprise task evaluation