# Prompt Caching in the Claude API with cache_control

> Cache a long system prompt or document with cache_control, confirm hits in the usage fields, and know when the 1.25x write cost pays for itself.

- Canonical: https://guides-ai.pages.dev/guides/claude-api-prompt-caching/
- Plate 15.04 · Topic: Building with the API (https://guides-ai.pages.dev/topics/api/)
- Published: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

If every request re-sends the same 40-page manual, you pay full input price every time. Prompt caching stores that prefix server-side and later requests read it at 0.1x the input rate.

## Mark the breakpoint

```python
import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

with open("handbook.md") as f:
    handbook = f.read()

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=1024,
    system=[
        {"type": "text", "text": "You answer questions about the handbook."},
        {"type": "text", "text": handbook, "cache_control": {"type": "ephemeral"}},
    ],
    messages=[{"role": "user", "content": "What is the refund window?"}],
)
```

Everything up to and including the marked block becomes the cached prefix; the varying question sits after it. Order matters — the API renders `tools`, then `system`, then `messages`.

```bash
curl https://api.anthropic.com/v1/messages \
  -H "content-type: application/json" \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "claude-opus-5",
    "max_tokens": 1024,
    "system": [
      {"type": "text", "text": "<the long document>", "cache_control": {"type": "ephemeral"}}
    ],
    "messages": [{"role": "user", "content": "Summarise section 3."}]
  }'
```

The same breakpoint over raw HTTP. A top-level `"cache_control": {"type": "ephemeral"}` on the request body auto-places the breakpoint on the last cacheable block if you would rather not annotate one by hand.

## Confirm it actually hit

```python
print(response.usage.cache_creation_input_tokens)  # tokens written to cache
print(response.usage.cache_read_input_tokens)      # tokens served from cache
print(response.usage.input_tokens)                 # everything after the breakpoint
```

The only honest check. If `cache_read_input_tokens` stays at 0 across repeated requests, something inside the prefix is changing: a `datetime.now()` in the system prompt, a request ID, an unsorted `json.dumps`, or a tool list built in nondeterministic order. Caching is a byte-exact prefix match — one changed character invalidates everything after it.

## When it pays off

- **Cost:** writes cost 1.25x base input (2x on the 1-hour TTL), reads 0.1x. Two hits and you are ahead.
- **TTL:** 5 minutes by default, refreshed on every hit. Pass `{"type": "ephemeral", "ttl": "1h"}` for an hour.
- **Minimum prefix:** model-dependent — 512 tokens on Claude Opus 5, 1,024 on Sonnet 5, 4,096 on Haiku 4.5. Shorter prefixes silently don't cache at all.
- **Limit:** at most 4 explicit breakpoints per request.

Chat histories, RAG context, long system prompts, and big tool definitions are the usual wins. A prompt that changes every call is not worth caching.

---

Next: [batch requests at half price](/guides/claude-api-batch-requests/) · [your first Claude API request](/guides/claude-api-first-request/).
