Model techniques map
Techniquesevaluationhuman & real-world evaluation

specific method · filed under evaluation

Blind side-by-side expert evaluation

Human experts compare or score model outputs without knowing which model produced them; the cited engineering evaluation involved internal experts and engineering tasks.

Also called blind expert judging of model outputs, blind expert judging, Blind side-by-side evaluation, Blind side-by-side human evaluation.

sources
4
models
2
labs adopt it
2
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 4

Documented in

Evidence

4 spans quoted from the sources, strongest treatment first.

we ran a blind side-by-side evaluation: 163 internal experts rated model outputs on 203 engineering tasks

usedevaluation onlyin Hy4-previewTencent Hunyuan

we ran a blind side-by-side evaluation: 163 internal experts rated model outputs on 203 engineering tasks.

usedevaluation onlyin Hy4-previewTencent Hunyuan

The comparison is performed under blind expert judging, where experts score each output on code quality, feature completeness, visual fidelity, and interaction experience without knowing which model produced it.

usedevaluation onlyin Kimi Webdev BenchMoonshot AI

In an internal blind evaluation involving 163 experts and 203 engineering tasks, Hy4 preview achieved an average score of 2.99 out of 4

usedevaluation onlyin Hy4-previewTencent

Filed alongside

Other methods under evaluation :: human & real-world evaluation.