Everyday tasks

How to Generate Speech Locally with Piper TTS

3 min read

Piper is a small neural text-to-speech engine that runs on CPU — fast enough on a Raspberry Pi, instant on a laptop. It is the simplest local TTS to install: one pip package, one voice download, one command. Nothing is uploaded.

1. Install

pip install piper-tts

Installs the piper module and its ONNX runtime. No GPU required. On Windows use python instead of python3 in every command below.

2. Download a voice

python3 -m piper.download_voices
python3 -m piper.download_voices en_US-lessac-medium

The first command lists every available voice; the second downloads one. Voices are named <language>_<REGION>-<name>-<quality>, where quality is x_low, low, medium or high — bigger is slower and better. Each voice is two files, an .onnx model and an .onnx.json config, saved into the current directory unless you pass --data-dir <DIR>.

Voices come from the rhasspy/piper-voices repository on Hugging Face and their licences differ — check the MODEL_CARD file next to a voice before you use it commercially.

3. Write a WAV

python3 -m piper -m en_US-lessac-medium -f test.wav -- 'This is a test.'

Writes test.wav. The -- matters: it ends the flags so text starting with a dash is not parsed as an option. Drop -f entirely and Piper plays the audio through ffplay instead of saving it.

4. Read from a file

python3 -m piper -m en_US-lessac-medium --input-file chapter1.txt -f chapter1.wav

Synthesizes a whole text file into one WAV. Useful extras: --sentence-silence 0.5 inserts half a second between sentences, --volume 1.5 boosts the level, and --cuda uses an NVIDIA GPU if you also installed onnxruntime-gpu.

Piper reloads the model on every invocation, so a loop over hundreds of short strings is slow — run its HTTP server instead for that.

5. If the voice is not good enough

Piper trades naturalness for speed. Kokoro is the other easy local option: an 82M-parameter model with noticeably better prosody, still CPU-friendly.

from kokoro import KPipeline
import soundfile as sf

pipeline = KPipeline(lang_code='a')          # 'a' = American English
for i, (_, _, audio) in enumerate(pipeline("Hello from a local model.", voice='af_heart')):
    sf.write(f"{i}.wav", audio, 24000)

Needs pip install kokoro soundfile plus the system package espeak-ng, and writes 24 kHz audio. All 54 voices ship inside the package.


Next: go the other direction and transcribe audio or summarize a video before narrating it.

Open the full interactive version (with copy buttons) ↗

← All guides