How to Transcribe Audio Locally with Whisper
Install OpenAI Whisper and ffmpeg, run whisper file.mp3 --model turbo, and get SRT and TXT subtitle files with timestamps out on your own machine.
On this page4 sections
Whisper transcribes audio offline — nothing is uploaded, and it handles about 100 languages. There are three common builds; this guide uses the reference openai-whisper package because it is one pip install and it writes SRT and TXT files itself. faster-whisper runs the same models roughly 4x quicker but is a Python library you have to format output with, and whisper.cpp is the choice when you want a single C++ binary and no Python at all.
1. Install
pip install -U openai-whisper
Installs the CLI and the Python package.
Whisper shells out to ffmpeg to decode audio, so install that too:
sudo apt update && sudo apt install ffmpeg # Debian/Ubuntu
brew install ffmpeg # macOS
choco install ffmpeg # or: scoop install ffmpeg
Puts ffmpeg on your PATH. Without it, every transcription fails at the decode step.
2. Transcribe
whisper interview.mp3 --model turbo --output_format srt --output_dir out
Transcribes the file and writes out/interview.srt with timestamps. Use --output_format txt for plain text, or omit the flag entirely to get every format at once (txt, vtt, srt, tsv, json).
The first run downloads the model into ~/.cache/whisper.
3. Pick a model size
| Model | Parameters | VRAM | Speed |
|---|---|---|---|
| tiny | 39M | ~1 GB | ~10x |
| base | 74M | ~1 GB | ~7x |
| small | 244M | ~2 GB | ~4x |
| medium | 769M | ~5 GB | ~2x |
| turbo | 809M | ~6 GB | ~8x |
| large | 1550M | ~10 GB | 1x |
turbo is the default and the best trade for most people — near large-v3 accuracy at a fraction of the time. Drop to small on a machine with little memory. No GPU is fine; it just runs on CPU and takes considerably longer.
4. Non-English audio
whisper podcast.m4a --language German --output_format srt
Skips language auto-detection, which is faster and avoids the occasional wrong guess on a quiet opening.
To translate speech into English instead of transcribing it, add --task translate and use a multilingual model — turbo is not trained for translation and will hand back the original language:
whisper podcast.m4a --model medium --language German --task translate
Produces English text from German audio.
Next: summarize a YouTube video with AI or run a local model to summarize the transcript privately.