Prompting

LLM-as-Judge Evals: Rubric and Verdict

4 min read

When “correct” can’t be a string match — summaries, explanations, support replies — you need a grader. A second model can be that grader, as long as you treat its verdicts as measurements with known biases, not as truth.

When not to use one

If an exact match, regex, or JSON schema check can decide it, use that. Judges cost a call per case and drift when you change the judge model. Reach for one only for open-ended text.

The rubric prompt

You are grading one answer against a rubric. Be strict. Assume nothing
that is not in the reference.

Criteria:
1. Every factual claim is supported by the reference.
2. It answers the question asked, not a nearby one.
3. It says it does not know rather than guessing when the reference is silent.

Return JSON only, no prose:
{"verdict": "pass" | "fail",
 "criterion_failed": 1 | 2 | 3 | null,
 "evidence": "at most 25 words quoted from the answer",
 "confidence": 0.0 to 1.0}

Three things make it work: criteria that can fail, a required quote so the judge has to point at real text, and a single field naming which criterion broke — that’s what turns a score into a bug report.

The loop

import json, anthropic

client = anthropic.Anthropic()

def judge(question, reference, answer):
    msg = client.messages.create(
        model="claude-opus-5",
        max_tokens=300,
        system=RUBRIC,                      # the text block above
        messages=[{"role": "user", "content":
                   f"<question>{question}</question>\n"
                   f"<reference>{reference}</reference>\n"
                   f"<answer>{answer}</answer>"}],
    )
    text = "".join(b.text for b in msg.content if b.type == "text")
    return json.loads(text)

failures, cases = [], load_eval_set()        # list of {question, reference}
for row in cases:
    v = judge(row["question"], row["reference"], run_my_app(row["question"]))
    if v["verdict"] == "fail":
        failures.append((row["question"], v["criterion_failed"], v["evidence"]))

print(f"{len(cases) - len(failures)}/{len(cases)} passed")
for q, c, ev in failures:
    print(f"  criterion {c}: {q}\n    -> {ev}")

Scores the whole set and prints why each failure failed. json.loads on raw text will eventually hit a stray sentence — harden it with structured output once the rubric stabilises.

Pitfalls that produce a flattering number


Next: evals with promptfoo · reduce hallucinations.

Open the full interactive version (with copy buttons) ↗

← All guides