Model techniques map
Techniquesevaluation

implementation detail · filed under evaluation

Standardized agent evaluation protocol

Specifying sampling counts, sampling parameters, context limits, and maximum agent steps for benchmark runs.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

All scaffolds use N=8 samples per task on DeepSWE v1.1 and N=3 on Terminal-Bench 2.1, with Linux containers, temperature=1.0, top_p=0.95, a 1M-token context limit, and max_steps=500 per agent.

usedevaluation onlyin DeepSeek-V4.1-FlashDeepSeek

Filed alongside

Other methods under evaluation.