Prompt Evals with promptfoo: One YAML, One Matrix
Install promptfoo, write a config with two prompts, three test cases and real assertions, run it, and read the pass/fail grid.
On this page5 sections
Comparing two prompts by eyeballing three outputs is not an eval. promptfoo turns “which prompt is better” into a grid you can re-run after every change.
1. Install
npx promptfoo@latest init
Scaffolds a promptfooconfig.yaml in the current folder. No global install needed; npm install -g promptfoo works too if you’d rather type promptfoo directly.
2. The config
description: Support reply — tone and next-step check
prompts:
- 'Answer the customer question in one sentence: {{question}}'
- |
You are a support agent. Answer in one sentence.
Do not apologise. End with one concrete next step.
Question: {{question}}
providers:
- anthropic:messages:claude-opus-5
- anthropic:messages:claude-haiku-4-5
defaultTest:
assert:
- type: latency
threshold: 8000
tests:
- vars:
question: 'My card was charged twice.'
assert:
- type: icontains
value: refund
- type: llm-rubric
value: 'Names one concrete next step. Does not promise a specific refund amount.'
provider: anthropic:messages:claude-opus-5
- vars:
question: 'How do I export my data?'
assert:
- type: contains
value: export
- vars:
question: 'Do you train on my data?'
assert:
- type: llm-rubric
value: 'Answers directly. Does not invent a policy URL or a date.'
provider: anthropic:messages:claude-opus-5
Two prompts × two providers × three cases = twelve runs. {{question}} is filled from each test’s vars. defaultTest applies to every case, so the latency ceiling is written once.
3. Run and view
npx promptfoo@latest eval
npx promptfoo@latest view
eval prints the grid in the terminal and caches results; view opens the same thing in a browser where you can click a cell to see the full output and which assertion failed.
4. Reading it
Each row is a test case, each column a prompt-and-provider pair, each cell pass/fail with the token cost. Look for the column that passes everywhere and look at the cheap model’s column — if claude-haiku-4-5 passes the same cases, prompt two just bought you a cost cut, not a quality win.
Assertions worth knowing
equals, contains, icontains and regex take value. is-json and contains-json validate structure, with an optional schema in value. cost and latency take threshold. javascript and python run your own check. llm-rubric and similar are the graded ones.
Two habits keep the numbers meaningful: pin the grading provider explicitly instead of inheriting a default, and write rubrics as things that can fail (“does not invent a URL”) rather than vague praise (“is helpful”). A rubric everything passes measures nothing.