# Prompt Evals with promptfoo: One YAML, One Matrix

> Install promptfoo, write a config with two prompts, three test cases and real assertions, run it, and read the pass/fail grid.

- Canonical: https://guides-ai.pages.dev/guides/evals-with-promptfoo/
- Plate 10.06 · Topic: Prompting (https://guides-ai.pages.dev/topics/prompting/)
- Published: 06 Sept 2026 · 4 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

Comparing two prompts by eyeballing three outputs is not an eval. promptfoo turns "which prompt is better" into a grid you can re-run after every change.

## 1. Install

```bash
npx promptfoo@latest init
```

Scaffolds a `promptfooconfig.yaml` in the current folder. No global install needed; `npm install -g promptfoo` works too if you'd rather type `promptfoo` directly.

## 2. The config

```yaml
description: Support reply — tone and next-step check

prompts:
  - 'Answer the customer question in one sentence: {{question}}'
  - |
    You are a support agent. Answer in one sentence.
    Do not apologise. End with one concrete next step.
    Question: {{question}}

providers:
  - anthropic:messages:claude-opus-5
  - anthropic:messages:claude-haiku-4-5

defaultTest:
  assert:
    - type: latency
      threshold: 8000

tests:
  - vars:
      question: 'My card was charged twice.'
    assert:
      - type: icontains
        value: refund
      - type: llm-rubric
        value: 'Names one concrete next step. Does not promise a specific refund amount.'
        provider: anthropic:messages:claude-opus-5
  - vars:
      question: 'How do I export my data?'
    assert:
      - type: contains
        value: export
  - vars:
      question: 'Do you train on my data?'
    assert:
      - type: llm-rubric
        value: 'Answers directly. Does not invent a policy URL or a date.'
        provider: anthropic:messages:claude-opus-5
```

Two prompts × two providers × three cases = twelve runs. `{{question}}` is filled from each test's `vars`. `defaultTest` applies to every case, so the latency ceiling is written once.

## 3. Run and view

```bash
npx promptfoo@latest eval
npx promptfoo@latest view
```

`eval` prints the grid in the terminal and caches results; `view` opens the same thing in a browser where you can click a cell to see the full output and which assertion failed.

## 4. Reading it

Each row is a test case, each column a prompt-and-provider pair, each cell pass/fail with the token cost. Look for the column that passes everywhere *and* look at the cheap model's column — if `claude-haiku-4-5` passes the same cases, prompt two just bought you a cost cut, not a quality win.

## Assertions worth knowing

`equals`, `contains`, `icontains` and `regex` take `value`. `is-json` and `contains-json` validate structure, with an optional schema in `value`. `cost` and `latency` take `threshold`. `javascript` and `python` run your own check. `llm-rubric` and `similar` are the graded ones.

Two habits keep the numbers meaningful: pin the grading `provider` explicitly instead of inheriting a default, and write rubrics as things that can *fail* ("does not invent a URL") rather than vague praise ("is helpful"). A rubric everything passes measures nothing.

---

Next: [LLM-as-judge evals](/guides/llm-as-judge-evals/) · [write better AI prompts](/guides/write-better-ai-prompts/).
