Open-Source Toolkit for Text-to-Speech Synthesis
TL;DR
- This article covers everything you need to know about open-source text-to-speech (TTS) toolkits. We look at what they are, how they works, and why their super-useful for video creators and other content producers on a budget. You'll discover some of the best options out there, and how to decide if open source is right for your project, or if you need something more.
The Coqui TTS open-source toolkit is a Python library for training and running text-to-speech models, including XTTS voice cloning. The original coqui-ai/TTS repo has had no release since December 2023. Today the maintained version is the Idiap fork, installed as coqui-tts, with release 0.27.5 out on January 26, 2026.
Last updated: October 6, 2026. GitHub, PyPI and Hugging Face pages were checked on that date.
This guide is for creators and developers who want to run speech synthesis on their own machine. It covers Coqui TTS's current status, its architecture, setup and the XTTS license. It also compares 4 other open-source speech synthesis libraries.
Key Takeaways
- Use the Idiap fork. The original coqui-ai/TTS repo's last release was v0.22.0 on December 12, 2023. The fork at idiap/coqui-ai-TTS calls the original "unmaintained" and ships as
coqui-ttson PyPI.- XTTS-v2 clones a voice from a 6-second clip in 17 languages, but its model license is non-commercial only.
- The toolkit code is MPL-2.0. Model weights can carry their own licenses, so check each model before you publish audio.
- Alternatives exist for every need: Piper for fast local voices, eSpeak NG for tiny devices and 100+ languages, ESPnet for research.
- Running models yourself takes setup. You'll need Python 3.10 or newer, PyTorch and some patience with dependencies.
On this page: Status · News · Architecture · Install · Voice cloning · Alternatives · Choosing · Hosted option · FAQ
What is the Coqui TTS open-source toolkit, and is it still maintained?
Coqui TTS is a deep learning toolkit for text-to-speech. Its GitHub description calls it "battle-tested in research and production" (coqui-ai/TTS, retrieved 2026-10-06). It bundles many model types, pretrained voices and training scripts in one Python package.
There are 2 GitHub repos, and the difference matters:
| Repo | PyPI package | Latest release | Status (October 6, 2026) |
|---|---|---|---|
| coqui-ai/TTS (original) | TTS | v0.22.0, December 12, 2023 | Not archived, but last code push was August 16, 2024 |
| idiap/coqui-ai-TTS (fork) | coqui-tts | v0.27.5, January 26, 2026 | Maintained; describes itself as a fork of the "original, unmaintained repository" |
Sources: GitHub API for coqui-ai/TTS, idiap/coqui-ai-TTS and PyPI coqui-tts, retrieved 2026-10-06.
Both repos use the Mozilla Public License 2.0 (MPL-2.0) for the code. The fork is maintained by the Idiap Research Institute team. If you follow an older tutorial that says pip install TTS, swap it for coqui-tts.
Coqui TTS news: recent releases
Searches for "Coqui TTS news" usually mean "what changed lately?" Here are the recent releases of the maintained fork, from its GitHub releases page (idiap/coqui-ai-TTS releases, retrieved 2026-10-06).
| Version | Date | Main change |
|---|---|---|
| v0.27.5 | January 26, 2026 | Fixed XTTS inference with newer Hugging Face Transformers |
| v0.27.4 | January 23, 2026 | Python 3.14 and PyTorch 2.10 support; PyTorch no longer installed by default |
| v0.27.3 | December 13, 2025 | Sentence-level timestamps; better PyTorch 2.9 support |
| v0.27.2 | September 25, 2025 | Python 3.13 support and documentation link fixes |
| v0.27.0 | July 14, 2025 | Speaker caching for cloned voices and a unified synthesize() interface |
There was no release in November 2025. The December 2025 release (v0.27.3) is the one most "late 2025" news refers to.
How Coqui TTS works: architecture in plain English
Most text-to-speech systems, Coqui included, turn text into audio in 3 stages. Think of it as a small assembly line.
- Text processing. The text is cleaned up and often turned into phonemes, the basic sounds of a language. "Dr." becomes "doctor", for example.
- Acoustic model. A neural network turns the phonemes into a spectrogram, a picture of how the sound's pitch and energy change over time. Coqui calls these "spectrogram models", such as Tacotron 2, Glow-TTS and FastSpeech 2.
- Vocoder. A second network turns the spectrogram into a real audio waveform. Coqui includes vocoders such as HiFi-GAN, MelGAN and WaveRNN.
Coqui also ships end-to-end models that do stages 2 and 3 in one network, such as VITS, YourTTS, XTTS, Tortoise and Bark (coqui-ai/TTS README, retrieved 2026-10-06). End-to-end models are simpler to run and often sound more natural.
Want a deeper look at each stage? Our explainer on neural network architectures for AI voice generation covers the acoustic models, and the guide to neural vocoder architectures covers stage 3.
How to install and run Coqui TTS
These steps follow the fork's README. They assume you're comfortable with a terminal.
- Check Python. The fork is tested with Python 3.10 up to (but not including) 3.15.
- Install PyTorch first. From v0.27.4, PyTorch is not included by default, so install it for your system using the official PyTorch instructions.
- Install the toolkit:
pip install coqui-tts - List the available models:
tts --list_models - Make your first file:
tts --text "Hello from my first open-source voice." --out_path hello.wav - Import it into your editor. Drop the WAV file into your video editor like any other audio clip.
You can also use it from Python. The README's single-speaker example loads a model and writes a file:
from TTS.api import TTS
tts = TTS("tts_models/de/thorsten/tacotron2-DDC")
tts.tts_to_file(text="Ich bin eine Testnachricht.", file_path="output.wav")
If installation fails, the cause is usually a Python or PyTorch version mismatch. Create a fresh virtual environment and match the versions in the README.
Coqui TTS voice cloning with XTTS-v2
XTTS-v2 is Coqui's best-known model. It can clone a voice from a short reference clip and speak in another language. Its Hugging Face model card says it needs "just a 6-second audio clip". It supports 17 languages, including English, Spanish, French, German, Arabic, Chinese, Japanese and Hindi (coqui/XTTS-v2, retrieved 2026-10-06).
In code, you pass a short WAV file of the target voice:
tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2")
wav = tts.tts(text="Hello world!", speaker_wav="my/cloning/audio.wav", language="en")
Check the license before you publish. The XTTS-v2 weights use the Coqui Public Model License 1.0.0. That license is for non-commercial use only. In its terms, non-commercial means you get no direct or indirect payment from the model or its output (XTTS-v2 license, retrieved 2026-10-06). So a monetized YouTube channel or a client video is likely outside it.
Only clone voices you have permission to use. For how this kind of cloning works under the hood, see our guide to zero-shot voice cloning.
Other open-source speech synthesis libraries compared
Coqui isn't the only choice. Here are 4 maintained or widely used options, with facts from each project's own GitHub page (retrieved 2026-10-06).
| Toolkit | Code license | Approach | Status | Best for |
|---|---|---|---|---|
| Coqui TTS (Idiap fork) | MPL-2.0 | Neural, many models, voice cloning via XTTS | Maintained | Trying many models and cloning in one package |
| Piper | GPL-3.0 (new repo); MIT (old repo) | Neural, "fast and local" | Old rhasspy/piper repo archived; development moved to OHF-Voice/piper1-gpl | Fast offline voices on modest hardware |
| eSpeak NG | GPL-3.0 | Formant synthesis, a few MB in size | Maintained | Tiny devices, screen readers, 100+ languages and accents |
| ESPnet | Apache 2.0 | Research toolkit with Tacotron 2, FastSpeech 2, VITS and more | Maintained | Researchers training and comparing models |
Sources: OHF-Voice/piper1-gpl, rhasspy/piper, eSpeak NG and ESPnet, retrieved 2026-10-06.
eSpeak NG sounds robotic next to neural models, but it's clear, small and covers many languages. Piper sits in the middle: neural quality with low resource needs. ESPnet is the most flexible, and the hardest to start with.
How to choose an open-source TTS toolkit
Match the toolkit to your goal, not to the longest feature list.
| If you want… | Start with | Why |
|---|---|---|
| To try many voices and models fast | Coqui TTS | Pretrained models and a one-line command |
| Voice cloning for personal projects | Coqui XTTS-v2 | 6-second clip, 17 languages, non-commercial license |
| Offline voices on a small computer | Piper | Built to be fast and local |
| A tiny engine with wide language coverage | eSpeak NG | A few MB, 100+ languages and accents |
| To train and publish research models | ESPnet | Recipes for many model types |
Also check 3 things before you commit:
- License of the code and of the model weights. They're often different.
- Recent activity. Look at the date of the last release, not just the star count.
- Your hardware. Large neural models run far faster on a GPU.
To compare output quality between toolkits, run the same script through each and score it with the listening tests in our guide to synthetic speech intelligibility metrics.
When a hosted AI voice generator makes more sense
Open-source toolkits are great for learning, privacy and custom research. They're less great when you just need a finished voiceover by this afternoon.
A hosted tool like Kveeky skips the setup: paste a script, pick a voice, download the audio. Kveeky has 700+ AI voices in 40+ languages and a free plan with 500 credits a month (about 6.6 minutes) and no credit card. Every paid plan includes commercial usage rights, which matters if XTTS's non-commercial license rules it out for you.
Where open source wins: you control the model, your audio never leaves your machine, and you can train on your own data. Where a hosted tool wins: no installs, no GPU, ready-made voices and clear commercial terms.
Frequently asked questions
Is Coqui TTS still maintained?
The original coqui-ai/TTS repo has had no release since v0.22.0 on December 12, 2023. The Idiap fork, installed as coqui-tts, is maintained. Its latest release was v0.27.5 on January 26, 2026.
Is Coqui TTS free for commercial use?
The toolkit code is MPL-2.0, which allows commercial use under its terms. Model weights have their own licenses. XTTS-v2 uses the Coqui Public Model License, which is non-commercial only, so check each model you use.
Where is the Coqui TTS GitHub repo?
The original is github.com/coqui-ai/TTS. The maintained fork is github.com/idiap/coqui-ai-TTS, and its PyPI package is coqui-tts.
Can Coqui TTS clone voices?
Yes. The XTTS-v2 model clones a voice from a short reference clip, about 6 seconds according to its model card, and supports 17 languages. Its license limits use to non-commercial projects.
What are the best open-source speech synthesis libraries?
Coqui TTS for many models and voice cloning, Piper for fast offline voices, eSpeak NG for tiny size and 100+ languages, and ESPnet for research. The best one depends on your hardware and license needs.
Do I need a GPU to run Coqui TTS?
Not always. Smaller models can run on a CPU, but large models like XTTS-v2 run much faster on a GPU. Lighter engines such as Piper or eSpeak NG are better fits for low-power machines.
How we checked this guide
This guide is written by Govind Kumar for the Kveeky team. Disclosure: Kveeky makes an AI voice generator. No Kveeky usage data is used in this guide.
- Coqui TTS repo status, license and release dates: GitHub coqui-ai/TTS, Idiap fork and PyPI coqui-tts, retrieved 2026-10-06.
- XTTS-v2 languages, clip length and license: Hugging Face model card and license text, retrieved 2026-10-06.
- Piper, eSpeak NG and ESPnet facts: each project's GitHub page, retrieved 2026-10-06.
- Kveeky plans: Kveeky pricing, retrieved 2026-10-06.
Your next step: if you're new to speech synthesis, read our plain-English guide to how text-to-speech AI works before you install anything. It'll make every model name above easier to follow.