Model techniques map
Techniquesevaluation

general family · filed under evaluation

Single-agent and multi-agent evaluation configurations

Comparing benchmark performance with single-agent and multi-agent configurations.

source
1
model
1
labs adopt it
0
strongest
evaluated

How sources treat it

One count per evidence span, weakest treatment to strongest.

evaluated 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

We evaluate single-agent and multi-agent configurations on both benchmarks under explicit per-rollout wall-clock deadlines.

evaluatedevaluation onlyin DeepSeek-V4.1-FlashDeepSeek

Filed alongside

Other methods under evaluation.