# Stream LLM Responses: Claude and OpenAI

> Print tokens as they arrive from the Claude and OpenAI APIs — Python and JavaScript, plus the raw SSE events and how to detect the end.

- Canonical: https://guides-ai.pages.dev/guides/stream-llm-responses/
- Plate 15.08 · Topic: Building with the API (https://guides-ai.pages.dev/topics/api/)
- Published: 06 Sept 2026 · 4 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

Streaming is not just a UX nicety: a non-streaming request with a large token limit can sit on an idle connection long enough to time out. Both APIs stream over Server-Sent Events.

## Claude — Python

```python
import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

with client.messages.stream(
    model="claude-opus-5",
    max_tokens=4000,
    messages=[{"role": "user", "content": "Write a haiku about rate limits."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

    final = stream.get_final_message()

print(f"\n\n[{final.usage.output_tokens} output tokens, {final.stop_reason}]")
```

`text_stream` yields plain text deltas; `get_final_message()` returns the accumulated `Message` once the stream closes, so you never have to reassemble it yourself. `flush=True` is what makes tokens actually appear.

## Claude — JavaScript

```javascript
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic(); // reads ANTHROPIC_API_KEY

const stream = client.messages.stream({
  model: "claude-opus-5",
  max_tokens: 4000,
  messages: [{ role: "user", content: "Write a haiku about rate limits." }],
});

for await (const event of stream) {
  if (event.type === "content_block_delta" && event.delta.type === "text_delta") {
    process.stdout.write(event.delta.text);
  }
}

const final = await stream.finalMessage();
console.log(`\n[${final.usage.output_tokens} output tokens]`);
```

Filters the event stream down to text deltas. Don't wrap `.on()` handlers in a `new Promise()` — `finalMessage()` already resolves on completion, error and abort.

## Claude — raw SSE

```bash
curl https://api.anthropic.com/v1/messages \
  -H "content-type: application/json" \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -d '{"model":"claude-opus-5","max_tokens":4000,"stream":true,
       "messages":[{"role":"user","content":"Write a haiku about rate limits."}]}'
```

Adding `"stream": true` turns the same endpoint into a Server-Sent Events feed. Each `data:` line is one JSON event:

```json
{"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}
```

The order is `message_start` → `content_block_start` → many `content_block_delta` → `content_block_stop` → `message_delta` (this one carries `stop_reason` and usage) → `message_stop`.

## OpenAI — Python

```python
from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY

stream = client.responses.create(
    model="gpt-6-astra",
    input="Write a haiku about rate limits.",
    stream=True,
)

for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
    elif event.type == "response.completed":
        print("\n[done]")
```

The Responses API emits **semantic** events rather than raw chunks: `response.created`, `response.output_text.delta`, `response.completed`, `error`, plus tool-specific ones. Branch on `event.type`.

## OpenAI — JavaScript

```javascript
import OpenAI from "openai";

const client = new OpenAI(); // reads OPENAI_API_KEY

const stream = await client.responses.create({
  model: "gpt-6-astra",
  input: "Write a haiku about rate limits.",
  stream: true,
});

for await (const event of stream) {
  if (event.type === "response.output_text.delta") process.stdout.write(event.delta);
  if (event.type === "response.completed") console.log("\n[done]");
}
```

Same events, async iterator. Both keys stay in the environment (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`) — never in the page you stream to.

## Handling the end

Treat "stream finished" and "answer finished" as different things. Claude's `stop_reason` of `max_tokens` means you truncated the reply; on both APIs a dropped connection leaves you with a partial answer, so buffer what you got before retrying rather than starting from an empty string.

---

Next: [your first Claude API request](/guides/claude-api-first-request/) · [your first OpenAI request](/guides/openai-api-first-request/).
