§15.05

Claude Message Batches API: Bulk Jobs at Half Price

Create a batch of Messages requests, poll processing_status, stream results by custom_id — and pay 50% less than synchronous calls.

published 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Building with the API Markdown

Step 5 of 6 · Build with the Claude API

On this page4 sections
  1. 1. Create the batch
  2. 2. Poll
  3. 3. Read the results
  4. Cost and limits

When you have thousands of items to classify, summarise or tag, and nobody is waiting on the answer, the Message Batches API runs them asynchronously at 50% of standard token prices.

1. Create the batch

import anthropic
from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
from anthropic.types.messages.batch_create_params import Request

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

reviews = ["Excellent quality!", "Terrible service, never again.", "It's okay."]

batch = client.messages.batches.create(
    requests=[
        Request(
            custom_id=f"review-{i}",
            params=MessageCreateParamsNonStreaming(
                model="claude-haiku-4-5",
                max_tokens=16,
                messages=[{
                    "role": "user",
                    "content": f"positive, negative or neutral — one word: {text}",
                }],
            ),
        )
        for i, text in enumerate(reviews)
    ]
)

print(batch.id, batch.processing_status)  # msgbatch_... in_progress

Submits three independent Messages requests as one job. custom_id must be unique within the batch and match ^[a-zA-Z0-9_-]{1,64}$.

curl https://api.anthropic.com/v1/messages/batches \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"requests": [
        {"custom_id": "review-0",
         "params": {"model": "claude-haiku-4-5", "max_tokens": 16,
                    "messages": [{"role": "user", "content": "One word sentiment: Excellent quality!"}]}}
      ]}'

The raw shape: each requests[].params is an ordinary Messages API body, so vision, tools and caching all work unchanged.

2. Poll

import time

while True:
    batch = client.messages.batches.retrieve(batch.id)
    if batch.processing_status == "ended":
        break
    time.sleep(30)

print(batch.request_counts)  # succeeded / errored / expired / canceled counts

processing_status goes in_progress → ended. Most batches finish within an hour; anything still running at 24 hours expires and is not billed.

3. Read the results

for result in client.messages.batches.results(batch.id):
    if result.result.type == "succeeded":
        text = next(b.text for b in result.result.message.content if b.type == "text")
        print(result.custom_id, "->", text)
    else:
        print(result.custom_id, "->", result.result.type)  # errored | canceled | expired

Streams the results file one record at a time. Results arrive in any order — key them by custom_id, never by list position.

Cost and limits

  • 50% off standard input and output token prices.
  • Up to 100,000 requests or 256 MB per batch, whichever comes first.
  • Results stay downloadable for 29 days after creation.
  • Prompt caching stacks with the discount: put the shared document in system with cache_control and the whole batch reuses it.

Batching plus a smaller model (claude-haiku-4-5 above) is usually the single biggest cost cut available on bulk work.


Next: prompt caching · your first Claude API request.

← All Building with the API plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list