PDFs go in a document content block, not an image one. Same three source types as images: base64, url, or a Files API file_id. No beta header — all active models support it.
1. Python: a local PDF
import base64, anthropic
client = anthropic.Anthropic()
with open("policy.pdf", "rb") as f:
pdf = base64.standard_b64encode(f.read()).decode()
msg = client.messages.create(
model="claude-opus-5",
max_tokens=2048,
messages=[{"role": "user", "content": [
{"type": "document",
"source": {"type": "base64",
"media_type": "application/pdf",
"data": pdf},
"title": "Policy v4",
"citations": {"enabled": True}},
{"type": "text", "text": "What is the cancellation window?"},
]}],
)
Sends the file and turns citations on. Put the document before the question, as with images.
2. Limits
| Limit | Value |
|---|---|
| Max request size | 32 MB (whole payload, everything included) |
| Max pages per request | 600 — or 100 when the request’s context window is under 1M tokens |
| Format | Standard PDF, no password or encryption |
Dense PDFs can exhaust the context window well before the page limit: each page costs roughly 1,500–3,000 text tokens plus image tokens, because every page is rendered to an image alongside its extracted text. That dual pass is what lets Claude answer questions about charts and scanned tables. For repeat use, upload once via the Files API and reference the file_id to keep payloads small.
3. Reading the citations back
for block in msg.content:
if block.type == "text":
print(block.text, end="")
for c in block.citations or []:
print(f" [p.{c.start_page_number}]", end="")
With citations on, the reply splits into several text blocks; the cited ones carry a citations list. For PDFs each entry is a page_location with cited_text, document_index, document_title, start_page_number (1-indexed) and end_page_number (exclusive). cited_text is free — it doesn’t count toward output tokens.
Two rules from the docs: citations must be enabled on all documents in a request or none of them, and they cannot be combined with structured outputs — sending output_config.format alongside an enabled document returns a 400.
4. When to pre-extract text instead
Send the PDF when the answer depends on layout: charts, diagrams, multi-column tables, forms, or scans with no text layer. Those are exactly the cases text extraction destroys.
Pre-extract with pypdf or pdfplumber and send plain text when the document is ordinary prose — a contract, a policy, a paper. You drop the per-page image tokens entirely, you control the chunking, and you can cache or index the text. Also pre-extract when you’re over 600 pages or 32 MB: split the file, or extract and retrieve the relevant pages first.