# LLM-as-Judge Evals: Rubric and Verdict

> Grade model outputs with a judge model: a copy-paste rubric prompt, a structured verdict, a Python scoring loop, and the biases that fake good scores.

- Canonical: https://guides-ai.pages.dev/guides/llm-as-judge-evals/
- Plate 10.07 · Topic: Prompting (https://guides-ai.pages.dev/topics/prompting/)
- Published: 06 Sept 2026 · 4 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

When "correct" can't be a string match — summaries, explanations, support replies — you need a grader. A second model can be that grader, as long as you treat its verdicts as measurements with known biases, not as truth.

## When *not* to use one

If an exact match, regex, or JSON schema check can decide it, use that. Judges cost a call per case and drift when you change the judge model. Reach for one only for open-ended text.

## The rubric prompt

```text
You are grading one answer against a rubric. Be strict. Assume nothing
that is not in the reference.

Criteria:
1. Every factual claim is supported by the reference.
2. It answers the question asked, not a nearby one.
3. It says it does not know rather than guessing when the reference is silent.

Return JSON only, no prose:
{"verdict": "pass" | "fail",
 "criterion_failed": 1 | 2 | 3 | null,
 "evidence": "at most 25 words quoted from the answer",
 "confidence": 0.0 to 1.0}
```

Three things make it work: criteria that can *fail*, a required quote so the judge has to point at real text, and a single field naming which criterion broke — that's what turns a score into a bug report.

## The loop

```python
import json, anthropic

client = anthropic.Anthropic()

def judge(question, reference, answer):
    msg = client.messages.create(
        model="claude-opus-5",
        max_tokens=300,
        system=RUBRIC,                      # the text block above
        messages=[{"role": "user", "content":
                   f"<question>{question}</question>\n"
                   f"<reference>{reference}</reference>\n"
                   f"<answer>{answer}</answer>"}],
    )
    text = "".join(b.text for b in msg.content if b.type == "text")
    return json.loads(text)

failures, cases = [], load_eval_set()        # list of {question, reference}
for row in cases:
    v = judge(row["question"], row["reference"], run_my_app(row["question"]))
    if v["verdict"] == "fail":
        failures.append((row["question"], v["criterion_failed"], v["evidence"]))

print(f"{len(cases) - len(failures)}/{len(cases)} passed")
for q, c, ev in failures:
    print(f"  criterion {c}: {q}\n    -> {ev}")
```

Scores the whole set and prints why each failure failed. `json.loads` on raw text will eventually hit a stray sentence — harden it with [structured output](/guides/claude-api-structured-output/) once the rubric stabilises.

## Pitfalls that produce a flattering number

- **Position bias.** In A/B comparisons, judges favour one slot. Run each pair in both orders and only count a win when both orders agree; disagreements are ties.
- **Self-preference.** Judges tend to rate text that resembles their own output higher. Grade with a different model than the one you're testing, or at minimum accept that a model grading itself is a weak signal.
- **Leniency drift.** A vague rubric passes everything. If your pass rate is above ~95% on the first run, the rubric is broken, not the app.
- **Uncalibrated judge.** Hand-label 30 cases yourself and check the judge against them. If it disagrees with you often, fix the rubric before you trust a single percentage.

---

Next: [evals with promptfoo](/guides/evals-with-promptfoo/) · [reduce hallucinations](/guides/reduce-ai-hallucinations/).
