When you have thousands of items to classify, summarise or tag, and nobody is waiting on the answer, the Message Batches API runs them asynchronously at 50% of standard token prices.
1. Create the batch
import anthropic
from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
from anthropic.types.messages.batch_create_params import Request
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
reviews = ["Excellent quality!", "Terrible service, never again.", "It's okay."]
batch = client.messages.batches.create(
requests=[
Request(
custom_id=f"review-{i}",
params=MessageCreateParamsNonStreaming(
model="claude-haiku-4-5",
max_tokens=16,
messages=[{
"role": "user",
"content": f"positive, negative or neutral — one word: {text}",
}],
),
)
for i, text in enumerate(reviews)
]
)
print(batch.id, batch.processing_status) # msgbatch_... in_progress
Submits three independent Messages requests as one job. custom_id must be unique within the batch and match ^[a-zA-Z0-9_-]{1,64}$.
curl https://api.anthropic.com/v1/messages/batches \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"requests": [
{"custom_id": "review-0",
"params": {"model": "claude-haiku-4-5", "max_tokens": 16,
"messages": [{"role": "user", "content": "One word sentiment: Excellent quality!"}]}}
]}'
The raw shape: each requests[].params is an ordinary Messages API body, so vision, tools and caching all work unchanged.
2. Poll
import time
while True:
batch = client.messages.batches.retrieve(batch.id)
if batch.processing_status == "ended":
break
time.sleep(30)
print(batch.request_counts) # succeeded / errored / expired / canceled counts
processing_status goes in_progress → ended. Most batches finish within an hour; anything still running at 24 hours expires and is not billed.
3. Read the results
for result in client.messages.batches.results(batch.id):
if result.result.type == "succeeded":
text = next(b.text for b in result.result.message.content if b.type == "text")
print(result.custom_id, "->", text)
else:
print(result.custom_id, "->", result.result.type) # errored | canceled | expired
Streams the results file one record at a time. Results arrive in any order — key them by custom_id, never by list position.
Cost and limits
- 50% off standard input and output token prices.
- Up to 100,000 requests or 256 MB per batch, whichever comes first.
- Results stay downloadable for 29 days after creation.
- Prompt caching stacks with the discount: put the shared document in
systemwithcache_controland the whole batch reuses it.
Batching plus a smaller model (claude-haiku-4-5 above) is usually the single biggest cost cut available on bulk work.