implementation detail · filed under evaluation
Behavior testing with stubs and edge cases
Testing generated scripts against stubbed dependencies and scenarios such as missing variables, real runs, and retention checks.
Also called behaviour test for every bash script.
- source
- 1
- models
- 2
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
scripts were run against stubbed pg_dump and aws binaries with a missing-variable test, a real run, and a retention check using fake 10-day and 3-day-old dumps
usedunclearin MiMo-V2.6-Pro and MiMo-V2.6-FlashXiaomi
Filed alongside
Other methods under evaluation.
Prompt-based output standardizationEvaluation without safety filtersRefusal-suppressed variants for capability estimationavg@kAlmost@1Deadline-bounded rollout evaluationDiverse benchmark validation for quantizationDocumenting evaluation configurationEnd-to-end exploit development evaluationEvaluation equivalence thresholdFixed MTP draft-token acceptance lengthFour-run mean pass@1Full trajectory releaseJoint validation protocolMean@5Model–harness co-designOfficial-release-source baseline restrictionPrompt engineering to elicit answersProtocolQA robustness validationRefusal behavior quantificationRollout-based auditingSample test and length-constraint filteringSandboxed GPU kernel optimization evaluationSingle-agent and multi-agent evaluation configurations