specific method · filed under evaluation
LLM-based anti-cheating judgment
LLM-based judgment is used to detect malicious behaviors such as unauthorized package or network operations.
Also called LLM-based judgement for anti-cheating, LLM-based judgement.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
To prevent hacking, we use rule-based and a LLM-based judgement to prevent malicious behaviors (e.g., unauthorized pip or curl operations).
usedunclearin GLM-5.3Z.ai
Filed alongside
Other methods under evaluation :: judge.
Agent-as-a-JudgeIndependent quality-inspection agentAgent-as-a-VerifierDecompositional instruction-following judgeDeterministic tool-call verifierGPT-5.5 (medium) judge modelHuman-calibrated LLM-as-a-JudgeLLM autograding with expert validationLLM-as-a-JudgeLLM-based correctness judgingLLM-based external API usage inspectionMandatory agentic judge protocolOfficial task verifier scoringReward-hack detection judgeRule-based anti-cheat checks