implementation detail · filed under other
Reward-Hacking Detection System
Detect and penalize kernel-task strategies such as CUDA graph replay, input caching, and precision reduction that exploit the reward.
Also called hacking-detection system.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
To ensure that rewards reflect genuine optimization, we develop a hacking-detection system that penalizes reward-hacking strategies such as CUDA graph replay, input caching, and precision reduction, and we continuously extend it with new safeguards as new hacking strategies are observed during Kimi K3’s development.
usedunclearin Kimi K3Moonshot AI
Filed alongside
Other methods under other.
Removing Git-History Leaks from Benchmark ImagesAgentic WorkflowsAutomating Repetitive Infrastructure WorkClosed-Model Performance with Refusal SubstitutionCode-Driven AnimationEarly Release with Feedback-Driven IterationGenerated Artifact Validation with Linters and Runtime TestsJoint Optimization of Architecture, Cache Precision, and DeploymentPrompt Addendum Prohibiting Direct Online-Solution UsePrompt Modality OrderingRemoving Pattern-Matching-Based Anti-Cheat ChecksRequest Caching in the Browser ToolScore-versus-Cost Efficiency ComparisonTraining–Inference ConsistencyUnique Run Identifiers