§15.12

Claude API Thinking and Effort

Turn adaptive thinking on, read the thinking blocks, dial depth with output_config.effort (low → max), stream deltas, and know what budget_tokens means.

published 06 Sept 2026 checked against docs 06 Sept 2026 4 min in Building with the API Markdown

On this page5 sections
  1. 1. Turn it on and make it visible
  2. 2. Size it with effort, not a token budget
  3. 3. Stream it
  4. 4. What it costs
  5. Rules that bite

Current Claude models reason before they answer. On the Claude 5 family (Opus 5, Sonnet 5, Fable 5.1) thinking is on by default but hidden; on Opus 4.6–4.8 and Sonnet 4.6 it’s off until you enable it. Either way, one parameter shows it and one parameter sizes it.

1. Turn it on and make it visible

import anthropic
client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=16000,
    thinking={"type": "adaptive", "display": "summarized"},
    messages=[{"role": "user", "content": "Are there infinitely many primes p with p mod 4 == 3? Prove it."}],
)
for block in response.content:
    if block.type == "thinking":
        print("THINKING:", block.thinking)
    elif block.type == "text":
        print("ANSWER:", block.text)

type: "adaptive" lets Claude decide whether and how much to think per request. display: "summarized" returns a readable summary of the reasoning; the default on Claude 5 models is "omitted", which returns thinking blocks with an empty thinking field (faster time-to-first-token — the encrypted signature still carries the reasoning for multi-turn continuity). No setting returns the raw chain of thought.

The same request in curl:

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"model":"claude-opus-5","max_tokens":16000,
       "thinking":{"type":"adaptive","display":"summarized"},
       "messages":[{"role":"user","content":"Explain why 0.1 + 0.2 != 0.3 in floating point."}]}'

Response content is an array: zero or more thinking blocks, then text. For a simple prompt at low effort, Claude may skip thinking entirely — no block appears.

2. Size it with effort, not a token budget

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=4096,
    output_config={"effort": "medium"},   # low · medium · high (default) · xhigh · max
    messages=[{"role": "user", "content": "Trade-offs of microservices vs a monolith, in 5 bullets."}],
)

effort scales the whole response — thinking, tool calls, prose. high equals not setting it; low is the cheap, fast setting for simple or high-volume work; xhigh/max are for long agentic and coding runs (raise max_tokens to 64k+ there — it is the hard ceiling for thinking plus answer). Changing top-level effort between requests invalidates prompt-cache prefixes, so pick a level per workload and keep it.

The older knob, thinking: {"type": "enabled", "budget_tokens": N}, still works on Claude 4.5 and earlier and is deprecated on 4.6 — Claude 4.7+ and every Claude 5 model reject it with a 400. Migrate: drop budget_tokens, use type: "adaptive", control depth with effort.

3. Stream it

with client.messages.stream(
    model="claude-opus-5", max_tokens=16000,
    thinking={"type": "adaptive", "display": "summarized"},
    messages=[{"role": "user", "content": "What is the GCD of 1071 and 462? Show the steps."}],
) as stream:
    for event in stream:
        if event.type == "content_block_delta":
            if event.delta.type == "thinking_delta":
                print(event.delta.thinking, end="", flush=True)
            elif event.delta.type == "text_delta":
                print(event.delta.text, end="", flush=True)

Thinking arrives as thinking_delta events inside content_block_delta, closed by one signature_delta, then text deltas follow. With display: "omitted" no thinking_delta is sent at all.

4. What it costs

Thinking tokens are billed as output tokens and count toward max_tokens even when hidden. Read the split in the response:

print(response.usage.output_tokens, response.usage.output_tokens_details.thinking_tokens)

When streaming, that breakdown appears only on the final message_delta event.

Rules that bite

  • With tools, pass thinking and redacted_thinking blocks back unchanged in the assistant turn; filtering on block.type == "thinking" alone silently breaks the loop.
  • On Claude 5 models (and Opus 4.7/4.8), non-default temperature, top_p or top_k return a 400 — leave them out.
  • thinking: {"type": "disabled"} works on Sonnet 5 and Opus 5 (except at xhigh/max effort); Fable 5.1 and Mythos models reject it.

Next: your first Claude API request · prompt caching · stream responses token by token.

← All Building with the API plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list