Model techniques map
Taxonomyevaluationjudge

taxonomy node · level 2

judge

16 methods filed at this node or below it, from the sources of 8 models.

evaluation :: judge

Matching aids for the classifier: LLM-as-a-judge; Agent-as-a-Judge; Agent-as-a-Verifier; judge model.

In this branch 16

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 16

Agent-as-a-Judge used · 1 source · 2 quotes
Agent-as-a-Verifier used · 1 source · 1 quote
Deterministic tool-call verifier used · 1 source · 1 quote
GPT-5.5 (medium) judge model used · 1 source · 1 quote
Human-calibrated LLM-as-a-Judge used · 1 source · 1 quote
Independent quality-inspection agent used · 1 source · 2 quotes
LLM autograding with expert validation used · 1 source · 1 quote
LLM-as-a-Judge used · 1 source · 1 quote
LLM-based anti-cheating judgment used · 1 source · 1 quote
LLM-based correctness judging used · 1 source · 1 quote
LLM-based external API usage inspection used · 1 source · 1 quote
Mandatory agentic judge protocol used · 1 source · 1 quote
Official task verifier scoring used · 1 source · 1 quote
Reward-hack detection judge used · 1 source · 1 quote
Rule-based anti-cheat checks used · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.