Everyday tasks

How to Translate an SRT Subtitle File with AI

4 min read

The mistake everyone makes first is pasting a whole .srt into a chat box. The model translates the timestamps too — shifting a digit here, merging two cues there — and the file no longer plays.

The fix is structural: never let the model see the timing lines. Parse the file yourself, send only the dialogue, rebuild from your own untouched indices and timestamps.

1. Install

pip install anthropic
export ANTHROPIC_API_KEY=sk-ant-...

Installs the SDK and puts your key where it is picked up. On Windows PowerShell: $env:ANTHROPIC_API_KEY = "sk-ant-...".

2. The script

import re, pathlib
from anthropic import Anthropic

SRC, DST, LANG, BATCH = "movie.srt", "movie.de.srt", "German", 40

raw = pathlib.Path(SRC).read_text(encoding="utf-8-sig").strip()
cues = []
for block in re.split(r"\n\s*\n", raw):
    lines = block.splitlines()
    if len(lines) >= 3:
        cues.append((lines[0], lines[1], " ".join(lines[2:])))   # index, timing, text

client = Anthropic()
SYSTEM = (
    f"You translate film subtitles into {LANG}. The input is numbered lines. "
    f"Return exactly one translated line per input line, with the same numbering and "
    f"nothing else: no notes, no blank lines, no merged or split lines. "
    f"Keep proper names, and keep HTML tags such as <i> exactly where they are."
)

out = []
for i in range(0, len(cues), BATCH):
    part = cues[i:i + BATCH]
    payload = "\n".join(f"{n}. {c[2]}" for n, c in enumerate(part, 1))
    msg = client.messages.create(
        model="claude-opus-5", max_tokens=16000, system=SYSTEM,
        messages=[{"role": "user", "content": payload}],
    )
    reply = "".join(b.text for b in msg.content if b.type == "text")
    got = [re.sub(r"^\d+\.\s*", "", ln) for ln in reply.strip().splitlines() if ln.strip()]
    assert len(got) == len(part), f"cue {i}: got {len(got)} lines, expected {len(part)}"
    out += [(c[0], c[1], t) for c, t in zip(part, got)]
    print(f"{len(out)}/{len(cues)}")

pathlib.Path(DST).write_text(
    "\n\n".join(f"{i}\n{t}\n{x}" for i, t, x in out) + "\n", encoding="utf-8")

Sends dialogue 40 cues at a time, keeps your original indices and timing strings, and asserts the line count on every batch so a drifting response fails loudly instead of corrupting the file. Multi-line cues are joined with a space — players re-wrap anyway, and one-line-in/one-line-out is what makes that check reliable. The reply is read by filtering content blocks on type == "text", not by grabbing content[0].

Batching is not optional: 40 cues per call give the model surrounding dialogue, so pronouns and formality stay consistent.

3. For a short file, prompt directly

Below are numbered subtitle lines. Translate them into German.
Return exactly one line per input line with the same numbering, nothing else.
Keep names and <i> tags. Keep each line short enough to read in the time a
subtitle allows — shorten rather than explain.

1. ...

Works in any chat UI — count the returned lines before rebuilding the file.

4. Check it

grep -c ' --> ' movie.srt movie.de.srt

Both files must report the same number of timing lines. Then open the translation in a player and skip to the last cue — if it is in sync there, the whole file is. Finally, scan for lines much longer than the original: nobody can read those in time.

Prefer local? Swap the client for the Ollama API with a multilingual model, keeping the batching and the assert. Quality on idiom drops.


Next: make the subtitles first with Whisper or learn the Claude API basics.

Open the full interactive version (with copy buttons) ↗

← All guides