Prompting

Prompt Evals with promptfoo: One YAML, One Matrix

4 min read

Comparing two prompts by eyeballing three outputs is not an eval. promptfoo turns “which prompt is better” into a grid you can re-run after every change.

1. Install

npx promptfoo@latest init

Scaffolds a promptfooconfig.yaml in the current folder. No global install needed; npm install -g promptfoo works too if you’d rather type promptfoo directly.

2. The config

description: Support reply — tone and next-step check

prompts:
  - 'Answer the customer question in one sentence: {{question}}'
  - |
    You are a support agent. Answer in one sentence.
    Do not apologise. End with one concrete next step.
    Question: {{question}}

providers:
  - anthropic:messages:claude-opus-5
  - anthropic:messages:claude-haiku-4-5

defaultTest:
  assert:
    - type: latency
      threshold: 8000

tests:
  - vars:
      question: 'My card was charged twice.'
    assert:
      - type: icontains
        value: refund
      - type: llm-rubric
        value: 'Names one concrete next step. Does not promise a specific refund amount.'
        provider: anthropic:messages:claude-opus-5
  - vars:
      question: 'How do I export my data?'
    assert:
      - type: contains
        value: export
  - vars:
      question: 'Do you train on my data?'
    assert:
      - type: llm-rubric
        value: 'Answers directly. Does not invent a policy URL or a date.'
        provider: anthropic:messages:claude-opus-5

Two prompts × two providers × three cases = twelve runs. {{question}} is filled from each test’s vars. defaultTest applies to every case, so the latency ceiling is written once.

3. Run and view

npx promptfoo@latest eval
npx promptfoo@latest view

eval prints the grid in the terminal and caches results; view opens the same thing in a browser where you can click a cell to see the full output and which assertion failed.

4. Reading it

Each row is a test case, each column a prompt-and-provider pair, each cell pass/fail with the token cost. Look for the column that passes everywhere and look at the cheap model’s column — if claude-haiku-4-5 passes the same cases, prompt two just bought you a cost cut, not a quality win.

Assertions worth knowing

equals, contains, icontains and regex take value. is-json and contains-json validate structure, with an optional schema in value. cost and latency take threshold. javascript and python run your own check. llm-rubric and similar are the graded ones.

Two habits keep the numbers meaningful: pin the grading provider explicitly instead of inheriting a default, and write rubrics as things that can fail (“does not invent a URL”) rather than vague praise (“is helpful”). A rubric everything passes measures nothing.


Next: LLM-as-judge evals · write better AI prompts.

Open the full interactive version (with copy buttons) ↗

← All guides