§10.06

Prompt Evals with promptfoo: One YAML, One Matrix

Install promptfoo, write a config with two prompts, three test cases and real assertions, run it, and read the pass/fail grid.

published 06 Sept 2026 checked against docs 06 Sept 2026 4 min in Prompting Markdown

On this page5 sections
  1. 1. Install
  2. 2. The config
  3. 3. Run and view
  4. 4. Reading it
  5. Assertions worth knowing

Comparing two prompts by eyeballing three outputs is not an eval. promptfoo turns “which prompt is better” into a grid you can re-run after every change.

1. Install

npx promptfoo@latest init

Scaffolds a promptfooconfig.yaml in the current folder. No global install needed; npm install -g promptfoo works too if you’d rather type promptfoo directly.

2. The config

description: Support reply — tone and next-step check

prompts:
  - 'Answer the customer question in one sentence: {{question}}'
  - |
    You are a support agent. Answer in one sentence.
    Do not apologise. End with one concrete next step.
    Question: {{question}}

providers:
  - anthropic:messages:claude-opus-5
  - anthropic:messages:claude-haiku-4-5

defaultTest:
  assert:
    - type: latency
      threshold: 8000

tests:
  - vars:
      question: 'My card was charged twice.'
    assert:
      - type: icontains
        value: refund
      - type: llm-rubric
        value: 'Names one concrete next step. Does not promise a specific refund amount.'
        provider: anthropic:messages:claude-opus-5
  - vars:
      question: 'How do I export my data?'
    assert:
      - type: contains
        value: export
  - vars:
      question: 'Do you train on my data?'
    assert:
      - type: llm-rubric
        value: 'Answers directly. Does not invent a policy URL or a date.'
        provider: anthropic:messages:claude-opus-5

Two prompts × two providers × three cases = twelve runs. {{question}} is filled from each test’s vars. defaultTest applies to every case, so the latency ceiling is written once.

3. Run and view

npx promptfoo@latest eval
npx promptfoo@latest view

eval prints the grid in the terminal and caches results; view opens the same thing in a browser where you can click a cell to see the full output and which assertion failed.

4. Reading it

Each row is a test case, each column a prompt-and-provider pair, each cell pass/fail with the token cost. Look for the column that passes everywhere and look at the cheap model’s column — if claude-haiku-4-5 passes the same cases, prompt two just bought you a cost cut, not a quality win.

Assertions worth knowing

equals, contains, icontains and regex take value. is-json and contains-json validate structure, with an optional schema in value. cost and latency take threshold. javascript and python run your own check. llm-rubric and similar are the graded ones.

Two habits keep the numbers meaningful: pin the grading provider explicitly instead of inheriting a default, and write rubrics as things that can fail (“does not invent a URL”) rather than vague praise (“is helpful”). A rubric everything passes measures nothing.


Next: LLM-as-judge evals · write better AI prompts.

← All Prompting plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list