§15.04

Prompt Caching in the Claude API with cache_control

Cache a long system prompt or document with cache_control, confirm hits in the usage fields, and know when the 1.25x write cost pays for itself.

published 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Building with the API Markdown

Step 4 of 6 · Build with the Claude API

On this page3 sections
  1. Mark the breakpoint
  2. Confirm it actually hit
  3. When it pays off

If every request re-sends the same 40-page manual, you pay full input price every time. Prompt caching stores that prefix server-side and later requests read it at 0.1x the input rate.

Mark the breakpoint

import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

with open("handbook.md") as f:
    handbook = f.read()

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=1024,
    system=[
        {"type": "text", "text": "You answer questions about the handbook."},
        {"type": "text", "text": handbook, "cache_control": {"type": "ephemeral"}},
    ],
    messages=[{"role": "user", "content": "What is the refund window?"}],
)

Everything up to and including the marked block becomes the cached prefix; the varying question sits after it. Order matters — the API renders tools, then system, then messages.

curl https://api.anthropic.com/v1/messages \
  -H "content-type: application/json" \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "claude-opus-5",
    "max_tokens": 1024,
    "system": [
      {"type": "text", "text": "<the long document>", "cache_control": {"type": "ephemeral"}}
    ],
    "messages": [{"role": "user", "content": "Summarise section 3."}]
  }'

The same breakpoint over raw HTTP. A top-level "cache_control": {"type": "ephemeral"} on the request body auto-places the breakpoint on the last cacheable block if you would rather not annotate one by hand.

Confirm it actually hit

print(response.usage.cache_creation_input_tokens)  # tokens written to cache
print(response.usage.cache_read_input_tokens)      # tokens served from cache
print(response.usage.input_tokens)                 # everything after the breakpoint

The only honest check. If cache_read_input_tokens stays at 0 across repeated requests, something inside the prefix is changing: a datetime.now() in the system prompt, a request ID, an unsorted json.dumps, or a tool list built in nondeterministic order. Caching is a byte-exact prefix match — one changed character invalidates everything after it.

When it pays off

  • Cost: writes cost 1.25x base input (2x on the 1-hour TTL), reads 0.1x. Two hits and you are ahead.
  • TTL: 5 minutes by default, refreshed on every hit. Pass {"type": "ephemeral", "ttl": "1h"} for an hour.
  • Minimum prefix: model-dependent — 512 tokens on Claude Opus 5, 1,024 on Sonnet 5, 4,096 on Haiku 4.5. Shorter prefixes silently don’t cache at all.
  • Limit: at most 4 explicit breakpoints per request.

Chat histories, RAG context, long system prompts, and big tool definitions are the usual wins. A prompt that changes every call is not worth caching.


Next: batch requests at half price · your first Claude API request.

← All Building with the API plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list