LLM-as-Judge Evals: Rubric and Verdict
Grade model outputs with a judge model: a copy-paste rubric prompt, a structured verdict, a Python scoring loop, and the biases that fake good scores.
On this page4 sections
When “correct” can’t be a string match — summaries, explanations, support replies — you need a grader. A second model can be that grader, as long as you treat its verdicts as measurements with known biases, not as truth.
When not to use one
If an exact match, regex, or JSON schema check can decide it, use that. Judges cost a call per case and drift when you change the judge model. Reach for one only for open-ended text.
The rubric prompt
You are grading one answer against a rubric. Be strict. Assume nothing
that is not in the reference.
Criteria:
1. Every factual claim is supported by the reference.
2. It answers the question asked, not a nearby one.
3. It says it does not know rather than guessing when the reference is silent.
Return JSON only, no prose:
{"verdict": "pass" | "fail",
"criterion_failed": 1 | 2 | 3 | null,
"evidence": "at most 25 words quoted from the answer",
"confidence": 0.0 to 1.0}
Three things make it work: criteria that can fail, a required quote so the judge has to point at real text, and a single field naming which criterion broke — that’s what turns a score into a bug report.
The loop
import json, anthropic
client = anthropic.Anthropic()
def judge(question, reference, answer):
msg = client.messages.create(
model="claude-opus-5",
max_tokens=300,
system=RUBRIC, # the text block above
messages=[{"role": "user", "content":
f"<question>{question}</question>\n"
f"<reference>{reference}</reference>\n"
f"<answer>{answer}</answer>"}],
)
text = "".join(b.text for b in msg.content if b.type == "text")
return json.loads(text)
failures, cases = [], load_eval_set() # list of {question, reference}
for row in cases:
v = judge(row["question"], row["reference"], run_my_app(row["question"]))
if v["verdict"] == "fail":
failures.append((row["question"], v["criterion_failed"], v["evidence"]))
print(f"{len(cases) - len(failures)}/{len(cases)} passed")
for q, c, ev in failures:
print(f" criterion {c}: {q}\n -> {ev}")
Scores the whole set and prints why each failure failed. json.loads on raw text will eventually hit a stray sentence — harden it with structured output once the rubric stabilises.
Pitfalls that produce a flattering number
- Position bias. In A/B comparisons, judges favour one slot. Run each pair in both orders and only count a win when both orders agree; disagreements are ties.
- Self-preference. Judges tend to rate text that resembles their own output higher. Grade with a different model than the one you’re testing, or at minimum accept that a model grading itself is a weak signal.
- Leniency drift. A vague rubric passes everything. If your pass rate is above ~95% on the first run, the rubric is broken, not the app.
- Uncalibrated judge. Hand-label 30 cases yourself and check the judge against them. If it disagrees with you often, fix the rubric before you trust a single percentage.