If every request re-sends the same 40-page manual, you pay full input price every time. Prompt caching stores that prefix server-side and later requests read it at 0.1x the input rate.
Mark the breakpoint
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
with open("handbook.md") as f:
handbook = f.read()
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
system=[
{"type": "text", "text": "You answer questions about the handbook."},
{"type": "text", "text": handbook, "cache_control": {"type": "ephemeral"}},
],
messages=[{"role": "user", "content": "What is the refund window?"}],
)
Everything up to and including the marked block becomes the cached prefix; the varying question sits after it. Order matters — the API renders tools, then system, then messages.
curl https://api.anthropic.com/v1/messages \
-H "content-type: application/json" \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-d '{
"model": "claude-opus-5",
"max_tokens": 1024,
"system": [
{"type": "text", "text": "<the long document>", "cache_control": {"type": "ephemeral"}}
],
"messages": [{"role": "user", "content": "Summarise section 3."}]
}'
The same breakpoint over raw HTTP. A top-level "cache_control": {"type": "ephemeral"} on the request body auto-places the breakpoint on the last cacheable block if you would rather not annotate one by hand.
Confirm it actually hit
print(response.usage.cache_creation_input_tokens) # tokens written to cache
print(response.usage.cache_read_input_tokens) # tokens served from cache
print(response.usage.input_tokens) # everything after the breakpoint
The only honest check. If cache_read_input_tokens stays at 0 across repeated requests, something inside the prefix is changing: a datetime.now() in the system prompt, a request ID, an unsorted json.dumps, or a tool list built in nondeterministic order. Caching is a byte-exact prefix match — one changed character invalidates everything after it.
When it pays off
- Cost: writes cost 1.25x base input (2x on the 1-hour TTL), reads 0.1x. Two hits and you are ahead.
- TTL: 5 minutes by default, refreshed on every hit. Pass
{"type": "ephemeral", "ttl": "1h"}for an hour. - Minimum prefix: model-dependent — 512 tokens on Claude Opus 5, 1,024 on Sonnet 5, 4,096 on Haiku 4.5. Shorter prefixes silently don’t cache at all.
- Limit: at most 4 explicit breakpoints per request.
Chat histories, RAG context, long system prompts, and big tool definitions are the usual wins. A prompt that changes every call is not worth caching.
Next: batch requests at half price · your first Claude API request.