specific method · filed under evaluation
Sandboxed GPU kernel optimization evaluation
Evaluating models on profiling, rewriting, and benchmarking tasks in identically configured sandboxes with up to 24 hours per task.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Each model works independently in an identically configured sandbox, with a budget of up to 24 hours per task for profiling, rewriting, and benchmarking.
usedevaluation onlyin GPU kernel optimization evaluation (AttnRes, DSA, KDA, MLA)Moonshot AI
Filed alongside
Other methods under evaluation.
Prompt-based output standardizationEvaluation without safety filtersRefusal-suppressed variants for capability estimationavg@kAlmost@1Behavior testing with stubs and edge casesDeadline-bounded rollout evaluationDiverse benchmark validation for quantizationDocumenting evaluation configurationEnd-to-end exploit development evaluationEvaluation equivalence thresholdFixed MTP draft-token acceptance lengthFour-run mean pass@1Full trajectory releaseJoint validation protocolMean@5Model–harness co-designOfficial-release-source baseline restrictionPrompt engineering to elicit answersProtocolQA robustness validationRefusal behavior quantificationRollout-based auditingSample test and length-constraint filteringSingle-agent and multi-agent evaluation configurations