Model techniques map
Techniquesevaluationhuman & real-world evaluation

specific method · filed under evaluation

Pairwise comparison for Chinese writing

Compares outputs pairwise on functional and creative Chinese writing tasks.

Also called Pairwise comparisons for Chinese writing.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

DeepSeek ran pairwise comparisons across 3,170 functional writing tasks and 2,837 creative writing tasks.

usedevaluation onlyin DeepSeek-V4DeepSeek

Filed alongside

Other methods under evaluation :: human & real-world evaluation.