specific method · filed under evaluation
mini-swe-agent harness
The mini-swe-agent harness is used to run DeepSWE evaluations.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
We run DeepSWE using the mini-swe-agent harness with temperature=0.95, top_p=1.0, timeout=6h and 400K context.
usedunclearin GLM-5.3Z.ai
Filed alongside
Other methods under evaluation :: evaluation harness.
DeepSeek Harness Minimal modeHarborNeMo Evaluator SDKAtomic binary rubric evaluationAutomatic in-flight checkpoint evaluation schedulingBenchmark evaluation frameworkCross-scaffold evaluationInline training evaluatorsIsolated container per rolloutJSON-structured answer-format promptNeMo Gym and NeMo Evaluator-based harnessNeMo SkillsPool