Current Claude models reason before they answer. On the Claude 5 family (Opus 5, Sonnet 5, Fable 5.1) thinking is on by default but hidden; on Opus 4.6–4.8 and Sonnet 4.6 it’s off until you enable it. Either way, one parameter shows it and one parameter sizes it.
1. Turn it on and make it visible
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
response = client.messages.create(
model="claude-opus-5",
max_tokens=16000,
thinking={"type": "adaptive", "display": "summarized"},
messages=[{"role": "user", "content": "Are there infinitely many primes p with p mod 4 == 3? Prove it."}],
)
for block in response.content:
if block.type == "thinking":
print("THINKING:", block.thinking)
elif block.type == "text":
print("ANSWER:", block.text)
type: "adaptive" lets Claude decide whether and how much to think per request. display: "summarized" returns a readable summary of the reasoning; the default on Claude 5 models is "omitted", which returns thinking blocks with an empty thinking field (faster time-to-first-token — the encrypted signature still carries the reasoning for multi-turn continuity). No setting returns the raw chain of thought.
The same request in curl:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model":"claude-opus-5","max_tokens":16000,
"thinking":{"type":"adaptive","display":"summarized"},
"messages":[{"role":"user","content":"Explain why 0.1 + 0.2 != 0.3 in floating point."}]}'
Response content is an array: zero or more thinking blocks, then text. For a simple prompt at low effort, Claude may skip thinking entirely — no block appears.
2. Size it with effort, not a token budget
response = client.messages.create(
model="claude-opus-5",
max_tokens=4096,
output_config={"effort": "medium"}, # low · medium · high (default) · xhigh · max
messages=[{"role": "user", "content": "Trade-offs of microservices vs a monolith, in 5 bullets."}],
)
effort scales the whole response — thinking, tool calls, prose. high equals not setting it; low is the cheap, fast setting for simple or high-volume work; xhigh/max are for long agentic and coding runs (raise max_tokens to 64k+ there — it is the hard ceiling for thinking plus answer). Changing top-level effort between requests invalidates prompt-cache prefixes, so pick a level per workload and keep it.
The older knob, thinking: {"type": "enabled", "budget_tokens": N}, still works on Claude 4.5 and earlier and is deprecated on 4.6 — Claude 4.7+ and every Claude 5 model reject it with a 400. Migrate: drop budget_tokens, use type: "adaptive", control depth with effort.
3. Stream it
with client.messages.stream(
model="claude-opus-5", max_tokens=16000,
thinking={"type": "adaptive", "display": "summarized"},
messages=[{"role": "user", "content": "What is the GCD of 1071 and 462? Show the steps."}],
) as stream:
for event in stream:
if event.type == "content_block_delta":
if event.delta.type == "thinking_delta":
print(event.delta.thinking, end="", flush=True)
elif event.delta.type == "text_delta":
print(event.delta.text, end="", flush=True)
Thinking arrives as thinking_delta events inside content_block_delta, closed by one signature_delta, then text deltas follow. With display: "omitted" no thinking_delta is sent at all.
4. What it costs
Thinking tokens are billed as output tokens and count toward max_tokens even when hidden. Read the split in the response:
print(response.usage.output_tokens, response.usage.output_tokens_details.thinking_tokens)
When streaming, that breakdown appears only on the final message_delta event.
Rules that bite
- With tools, pass
thinkingandredacted_thinkingblocks back unchanged in the assistant turn; filtering onblock.type == "thinking"alone silently breaks the loop. - On Claude 5 models (and Opus 4.7/4.8), non-default
temperature,top_portop_kreturn a 400 — leave them out. thinking: {"type": "disabled"}works on Sonnet 5 and Opus 5 (except atxhigh/maxeffort); Fable 5.1 and Mythos models reject it.
Next: your first Claude API request · prompt caching · stream responses token by token.