Building with the API

Stream LLM Responses: Claude and OpenAI

4 min read

Streaming is not just a UX nicety: a non-streaming request with a large token limit can sit on an idle connection long enough to time out. Both APIs stream over Server-Sent Events.

Claude — Python

import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

with client.messages.stream(
    model="claude-opus-5",
    max_tokens=4000,
    messages=[{"role": "user", "content": "Write a haiku about rate limits."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

    final = stream.get_final_message()

print(f"\n\n[{final.usage.output_tokens} output tokens, {final.stop_reason}]")

text_stream yields plain text deltas; get_final_message() returns the accumulated Message once the stream closes, so you never have to reassemble it yourself. flush=True is what makes tokens actually appear.

Claude — JavaScript

import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic(); // reads ANTHROPIC_API_KEY

const stream = client.messages.stream({
  model: "claude-opus-5",
  max_tokens: 4000,
  messages: [{ role: "user", content: "Write a haiku about rate limits." }],
});

for await (const event of stream) {
  if (event.type === "content_block_delta" && event.delta.type === "text_delta") {
    process.stdout.write(event.delta.text);
  }
}

const final = await stream.finalMessage();
console.log(`\n[${final.usage.output_tokens} output tokens]`);

Filters the event stream down to text deltas. Don’t wrap .on() handlers in a new Promise() — finalMessage() already resolves on completion, error and abort.

Claude — raw SSE

curl https://api.anthropic.com/v1/messages \
  -H "content-type: application/json" \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -d '{"model":"claude-opus-5","max_tokens":4000,"stream":true,
       "messages":[{"role":"user","content":"Write a haiku about rate limits."}]}'

Adding "stream": true turns the same endpoint into a Server-Sent Events feed. Each data: line is one JSON event:

{"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}

The order is message_start → content_block_start → many content_block_delta → content_block_stop → message_delta (this one carries stop_reason and usage) → message_stop.

OpenAI — Python

from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY

stream = client.responses.create(
    model="gpt-6-astra",
    input="Write a haiku about rate limits.",
    stream=True,
)

for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
    elif event.type == "response.completed":
        print("\n[done]")

The Responses API emits semantic events rather than raw chunks: response.created, response.output_text.delta, response.completed, error, plus tool-specific ones. Branch on event.type.

OpenAI — JavaScript

import OpenAI from "openai";

const client = new OpenAI(); // reads OPENAI_API_KEY

const stream = await client.responses.create({
  model: "gpt-6-astra",
  input: "Write a haiku about rate limits.",
  stream: true,
});

for await (const event of stream) {
  if (event.type === "response.output_text.delta") process.stdout.write(event.delta);
  if (event.type === "response.completed") console.log("\n[done]");
}

Same events, async iterator. Both keys stay in the environment (ANTHROPIC_API_KEY, OPENAI_API_KEY) — never in the page you stream to.

Handling the end

Treat “stream finished” and “answer finished” as different things. Claude’s stop_reason of max_tokens means you truncated the reply; on both APIs a dropped connection leaves you with a partial answer, so buffer what you got before retrying rather than starting from an empty string.


Next: your first Claude API request · your first OpenAI request.

Open the full interactive version (with copy buttons) ↗

← All guides