§11.04

OpenAI API vs Claude API: Which to Build On

OpenAI and Claude API compared on price per million tokens, 1M-token context, tool use and long-context surcharges — plus how to decide with your own eval.

published 02 Sept 2026 updated 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Choosing tools Markdown

On this page4 sections
  1. Price and size, side by side
  2. The costs that don’t show up in the table
  3. Tool use and structure
  4. Decide with an eval, not a table

Neither API is generically better, and the gap that used to decide it — context window — has closed. Both vendors’ current flagships hold about a million tokens. What is left is price per tier, how your workload maps onto that tier, and which SDK your codebase already speaks.

Price and size, side by side

Published list prices per million tokens, September 2026:

ModelContextInputOutput
GPT-6 Astra1.05M$10$50
Claude Opus 51M$5$25
GPT-5.6 Sol1.05M$4$20
GPT-5.6 Terra1.05M$2$12
Claude Sonnet 51M$2$10
Claude Haiku 4.5200K$1$5

Compare across rows, not vendors: Sonnet 5 and Terra are the same tier at nearly the same price, and either is the sane default. Reach for a top row only when an eval says the cheaper tier fails. Model naming is a moving target — pin the exact model ID rather than a floating alias.

The costs that don’t show up in the table

Long inputs can change the rate. OpenAI charges 2× input and 1.5× output on requests over 272K input tokens for the whole session, so a “cheap” model streaming a million tokens is not cheap. Discounts run the other way: Anthropic’s Batch API is 50% off and cache reads cost 10% of base input, which reshapes any workload with a stable prefix — see prompt caching and batch requests.

Tool use and structure

Both APIs define tools as JSON schemas, let the model choose when to call one, and loop until the task is done; both enforce structured output against a schema. The mechanics differ enough to matter when porting — compare Claude tool use with OpenAI structured outputs — but no idea in one is missing from the other.

Decide with an eval, not a table

Run this against both, on twenty real inputs, before you commit:

You are scoring two model outputs for the same task.
Task spec: <paste your actual prompt>
Output A: <model 1>
Output B: <model 2>
Score each 1-5 on: instruction-following, factual accuracy, format
compliance, and verbosity. Return a markdown table, then one sentence
naming the cheaper output that still passes.

Twenty of your own inputs beat any benchmark. Then read the numbers off the Claude models overview, because prices move.

← All Choosing tools plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list