specific method · filed under evaluation
Blind side-by-side expert evaluation
Human experts compare or score model outputs without knowing which model produced them; the cited engineering evaluation involved internal experts and engineering tasks.
Also called blind expert judging of model outputs, blind expert judging, Blind side-by-side evaluation, Blind side-by-side human evaluation.
- sources
- 4
- models
- 2
- labs adopt it
- 2
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
Documented in
Evidence
4 spans quoted from the sources, strongest treatment first.
we ran a blind side-by-side evaluation: 163 internal experts rated model outputs on 203 engineering tasks
we ran a blind side-by-side evaluation: 163 internal experts rated model outputs on 203 engineering tasks.
The comparison is performed under blind expert judging, where experts score each output on code quality, feature completeness, visual fidelity, and interaction experience without knowing which model produced it.
In an internal blind evaluation involving 163 experts and 203 engineering tasks, Hy4 preview achieved an average score of 2.99 out of 4
Filed alongside
Other methods under evaluation :: human & real-world evaluation.