Stream LLM Responses: Claude and OpenAI
Print tokens as they arrive from the Claude and OpenAI APIs — Python and JavaScript, plus the raw SSE events and how to detect the end.
Step 6 of 6 · Build with the Claude API
On this page6 sections
Streaming is not just a UX nicety: a non-streaming request with a large token limit can sit on an idle connection long enough to time out. Both APIs stream over Server-Sent Events.
Claude — Python
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
with client.messages.stream(
model="claude-opus-5",
max_tokens=4000,
messages=[{"role": "user", "content": "Write a haiku about rate limits."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
final = stream.get_final_message()
print(f"\n\n[{final.usage.output_tokens} output tokens, {final.stop_reason}]")
text_stream yields plain text deltas; get_final_message() returns the accumulated Message once the stream closes, so you never have to reassemble it yourself. flush=True is what makes tokens actually appear.
Claude — JavaScript
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic(); // reads ANTHROPIC_API_KEY
const stream = client.messages.stream({
model: "claude-opus-5",
max_tokens: 4000,
messages: [{ role: "user", content: "Write a haiku about rate limits." }],
});
for await (const event of stream) {
if (event.type === "content_block_delta" && event.delta.type === "text_delta") {
process.stdout.write(event.delta.text);
}
}
const final = await stream.finalMessage();
console.log(`\n[${final.usage.output_tokens} output tokens]`);
Filters the event stream down to text deltas. Don’t wrap .on() handlers in a new Promise() — finalMessage() already resolves on completion, error and abort.
Claude — raw SSE
curl https://api.anthropic.com/v1/messages \
-H "content-type: application/json" \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-d '{"model":"claude-opus-5","max_tokens":4000,"stream":true,
"messages":[{"role":"user","content":"Write a haiku about rate limits."}]}'
Adding "stream": true turns the same endpoint into a Server-Sent Events feed. Each data: line is one JSON event:
{"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}
The order is message_start → content_block_start → many content_block_delta → content_block_stop → message_delta (this one carries stop_reason and usage) → message_stop.
OpenAI — Python
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY
stream = client.responses.create(
model="gpt-6-astra",
input="Write a haiku about rate limits.",
stream=True,
)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
elif event.type == "response.completed":
print("\n[done]")
The Responses API emits semantic events rather than raw chunks: response.created, response.output_text.delta, response.completed, error, plus tool-specific ones. Branch on event.type.
OpenAI — JavaScript
import OpenAI from "openai";
const client = new OpenAI(); // reads OPENAI_API_KEY
const stream = await client.responses.create({
model: "gpt-6-astra",
input: "Write a haiku about rate limits.",
stream: true,
});
for await (const event of stream) {
if (event.type === "response.output_text.delta") process.stdout.write(event.delta);
if (event.type === "response.completed") console.log("\n[done]");
}
Same events, async iterator. Both keys stay in the environment (ANTHROPIC_API_KEY, OPENAI_API_KEY) — never in the page you stream to.
Handling the end
Treat “stream finished” and “answer finished” as different things. Claude’s stop_reason of max_tokens means you truncated the reply; on both APIs a dropped connection leaves you with a partial answer, so buffer what you got before retrying rather than starting from an empty string.
Next: your first Claude API request · your first OpenAI request.