§10.07

LLM-as-Judge Evals: Rubric and Verdict

Grade model outputs with a judge model: a copy-paste rubric prompt, a structured verdict, a Python scoring loop, and the biases that fake good scores.

published 06 Sept 2026 checked against docs 06 Sept 2026 4 min in Prompting Markdown

On this page4 sections
  1. When not to use one
  2. The rubric prompt
  3. The loop
  4. Pitfalls that produce a flattering number

When “correct” can’t be a string match — summaries, explanations, support replies — you need a grader. A second model can be that grader, as long as you treat its verdicts as measurements with known biases, not as truth.

When not to use one

If an exact match, regex, or JSON schema check can decide it, use that. Judges cost a call per case and drift when you change the judge model. Reach for one only for open-ended text.

The rubric prompt

You are grading one answer against a rubric. Be strict. Assume nothing
that is not in the reference.

Criteria:
1. Every factual claim is supported by the reference.
2. It answers the question asked, not a nearby one.
3. It says it does not know rather than guessing when the reference is silent.

Return JSON only, no prose:
{"verdict": "pass" | "fail",
 "criterion_failed": 1 | 2 | 3 | null,
 "evidence": "at most 25 words quoted from the answer",
 "confidence": 0.0 to 1.0}

Three things make it work: criteria that can fail, a required quote so the judge has to point at real text, and a single field naming which criterion broke — that’s what turns a score into a bug report.

The loop

import json, anthropic

client = anthropic.Anthropic()

def judge(question, reference, answer):
    msg = client.messages.create(
        model="claude-opus-5",
        max_tokens=300,
        system=RUBRIC,                      # the text block above
        messages=[{"role": "user", "content":
                   f"<question>{question}</question>\n"
                   f"<reference>{reference}</reference>\n"
                   f"<answer>{answer}</answer>"}],
    )
    text = "".join(b.text for b in msg.content if b.type == "text")
    return json.loads(text)

failures, cases = [], load_eval_set()        # list of {question, reference}
for row in cases:
    v = judge(row["question"], row["reference"], run_my_app(row["question"]))
    if v["verdict"] == "fail":
        failures.append((row["question"], v["criterion_failed"], v["evidence"]))

print(f"{len(cases) - len(failures)}/{len(cases)} passed")
for q, c, ev in failures:
    print(f"  criterion {c}: {q}\n    -> {ev}")

Scores the whole set and prints why each failure failed. json.loads on raw text will eventually hit a stray sentence — harden it with structured output once the rubric stabilises.

Pitfalls that produce a flattering number

  • Position bias. In A/B comparisons, judges favour one slot. Run each pair in both orders and only count a win when both orders agree; disagreements are ties.
  • Self-preference. Judges tend to rate text that resembles their own output higher. Grade with a different model than the one you’re testing, or at minimum accept that a model grading itself is a weak signal.
  • Leniency drift. A vague rubric passes everything. If your pass rate is above ~95% on the first run, the rubric is broken, not the app.
  • Uncalibrated judge. Hand-label 30 cases yourself and check the judge against them. If it disagrees with you often, fix the rubric before you trust a single percentage.

Next: evals with promptfoo · reduce hallucinations.

← All Prompting plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list