Open-Source Toolkit for Text-to-Speech Synthesis

Coqui TTS open-source toolkit Coqui TTS GitHub open-source speech synthesis libraries
Govind Kumar
Govind Kumar

Co-Founder & CTPO

 
September 14, 2025
9 min read
Open-Source Toolkit for Text-to-Speech Synthesis

TL;DR

  • This article covers everything you need to know about open-source text-to-speech (TTS) toolkits. We look at what they are, how they works, and why their super-useful for video creators and other content producers on a budget. You'll discover some of the best options out there, and how to decide if open source is right for your project, or if you need something more.

The Coqui TTS open-source toolkit is a Python library for training and running text-to-speech models, including XTTS voice cloning. The original coqui-ai/TTS repo has had no release since December 2023. Today the maintained version is the Idiap fork, installed as coqui-tts, with release 0.27.5 out on January 26, 2026.

Last updated: October 6, 2026. GitHub, PyPI and Hugging Face pages were checked on that date.

This guide is for creators and developers who want to run speech synthesis on their own machine. It covers Coqui TTS's current status, its architecture, setup and the XTTS license. It also compares 4 other open-source speech synthesis libraries.

Key Takeaways

  • Use the Idiap fork. The original coqui-ai/TTS repo's last release was v0.22.0 on December 12, 2023. The fork at idiap/coqui-ai-TTS calls the original "unmaintained" and ships as coqui-tts on PyPI.
  • XTTS-v2 clones a voice from a 6-second clip in 17 languages, but its model license is non-commercial only.
  • The toolkit code is MPL-2.0. Model weights can carry their own licenses, so check each model before you publish audio.
  • Alternatives exist for every need: Piper for fast local voices, eSpeak NG for tiny devices and 100+ languages, ESPnet for research.
  • Running models yourself takes setup. You'll need Python 3.10 or newer, PyTorch and some patience with dependencies.

On this page: Status · News · Architecture · Install · Voice cloning · Alternatives · Choosing · Hosted option · FAQ

What is the Coqui TTS open-source toolkit, and is it still maintained?

Coqui TTS is a deep learning toolkit for text-to-speech. Its GitHub description calls it "battle-tested in research and production" (coqui-ai/TTS, retrieved 2026-10-06). It bundles many model types, pretrained voices and training scripts in one Python package.

There are 2 GitHub repos, and the difference matters:

RepoPyPI packageLatest releaseStatus (October 6, 2026)
coqui-ai/TTS (original)TTSv0.22.0, December 12, 2023Not archived, but last code push was August 16, 2024
idiap/coqui-ai-TTS (fork)coqui-ttsv0.27.5, January 26, 2026Maintained; describes itself as a fork of the "original, unmaintained repository"

Sources: GitHub API for coqui-ai/TTS, idiap/coqui-ai-TTS and PyPI coqui-tts, retrieved 2026-10-06.

Both repos use the Mozilla Public License 2.0 (MPL-2.0) for the code. The fork is maintained by the Idiap Research Institute team. If you follow an older tutorial that says pip install TTS, swap it for coqui-tts.

Coqui TTS news: recent releases

Searches for "Coqui TTS news" usually mean "what changed lately?" Here are the recent releases of the maintained fork, from its GitHub releases page (idiap/coqui-ai-TTS releases, retrieved 2026-10-06).

VersionDateMain change
v0.27.5January 26, 2026Fixed XTTS inference with newer Hugging Face Transformers
v0.27.4January 23, 2026Python 3.14 and PyTorch 2.10 support; PyTorch no longer installed by default
v0.27.3December 13, 2025Sentence-level timestamps; better PyTorch 2.9 support
v0.27.2September 25, 2025Python 3.13 support and documentation link fixes
v0.27.0July 14, 2025Speaker caching for cloned voices and a unified synthesize() interface

There was no release in November 2025. The December 2025 release (v0.27.3) is the one most "late 2025" news refers to.

How Coqui TTS works: architecture in plain English

Most text-to-speech systems, Coqui included, turn text into audio in 3 stages. Think of it as a small assembly line.

  1. Text processing. The text is cleaned up and often turned into phonemes, the basic sounds of a language. "Dr." becomes "doctor", for example.
  2. Acoustic model. A neural network turns the phonemes into a spectrogram, a picture of how the sound's pitch and energy change over time. Coqui calls these "spectrogram models", such as Tacotron 2, Glow-TTS and FastSpeech 2.
  3. Vocoder. A second network turns the spectrogram into a real audio waveform. Coqui includes vocoders such as HiFi-GAN, MelGAN and WaveRNN.

Coqui also ships end-to-end models that do stages 2 and 3 in one network, such as VITS, YourTTS, XTTS, Tortoise and Bark (coqui-ai/TTS README, retrieved 2026-10-06). End-to-end models are simpler to run and often sound more natural.

Want a deeper look at each stage? Our explainer on neural network architectures for AI voice generation covers the acoustic models, and the guide to neural vocoder architectures covers stage 3.

How to install and run Coqui TTS

These steps follow the fork's README. They assume you're comfortable with a terminal.

  1. Check Python. The fork is tested with Python 3.10 up to (but not including) 3.15.
  2. Install PyTorch first. From v0.27.4, PyTorch is not included by default, so install it for your system using the official PyTorch instructions.
  3. Install the toolkit: pip install coqui-tts
  4. List the available models: tts --list_models
  5. Make your first file: tts --text "Hello from my first open-source voice." --out_path hello.wav
  6. Import it into your editor. Drop the WAV file into your video editor like any other audio clip.

You can also use it from Python. The README's single-speaker example loads a model and writes a file:

from TTS.api import TTS

tts = TTS("tts_models/de/thorsten/tacotron2-DDC") tts.tts_to_file(text="Ich bin eine Testnachricht.", file_path="output.wav")

If installation fails, the cause is usually a Python or PyTorch version mismatch. Create a fresh virtual environment and match the versions in the README.

Coqui TTS voice cloning with XTTS-v2

XTTS-v2 is Coqui's best-known model. It can clone a voice from a short reference clip and speak in another language. Its Hugging Face model card says it needs "just a 6-second audio clip". It supports 17 languages, including English, Spanish, French, German, Arabic, Chinese, Japanese and Hindi (coqui/XTTS-v2, retrieved 2026-10-06).

In code, you pass a short WAV file of the target voice:

tts = TTS("tts_models/multilingual/multi-dataset/xtts_v2")
wav = tts.tts(text="Hello world!", speaker_wav="my/cloning/audio.wav", language="en")

Check the license before you publish. The XTTS-v2 weights use the Coqui Public Model License 1.0.0. That license is for non-commercial use only. In its terms, non-commercial means you get no direct or indirect payment from the model or its output (XTTS-v2 license, retrieved 2026-10-06). So a monetized YouTube channel or a client video is likely outside it.

Only clone voices you have permission to use. For how this kind of cloning works under the hood, see our guide to zero-shot voice cloning.

Other open-source speech synthesis libraries compared

Coqui isn't the only choice. Here are 4 maintained or widely used options, with facts from each project's own GitHub page (retrieved 2026-10-06).

ToolkitCode licenseApproachStatusBest for
Coqui TTS (Idiap fork)MPL-2.0Neural, many models, voice cloning via XTTSMaintainedTrying many models and cloning in one package
PiperGPL-3.0 (new repo); MIT (old repo)Neural, "fast and local"Old rhasspy/piper repo archived; development moved to OHF-Voice/piper1-gplFast offline voices on modest hardware
eSpeak NGGPL-3.0Formant synthesis, a few MB in sizeMaintainedTiny devices, screen readers, 100+ languages and accents
ESPnetApache 2.0Research toolkit with Tacotron 2, FastSpeech 2, VITS and moreMaintainedResearchers training and comparing models

Sources: OHF-Voice/piper1-gpl, rhasspy/piper, eSpeak NG and ESPnet, retrieved 2026-10-06.

eSpeak NG sounds robotic next to neural models, but it's clear, small and covers many languages. Piper sits in the middle: neural quality with low resource needs. ESPnet is the most flexible, and the hardest to start with.

How to choose an open-source TTS toolkit

Match the toolkit to your goal, not to the longest feature list.

If you want…Start withWhy
To try many voices and models fastCoqui TTSPretrained models and a one-line command
Voice cloning for personal projectsCoqui XTTS-v26-second clip, 17 languages, non-commercial license
Offline voices on a small computerPiperBuilt to be fast and local
A tiny engine with wide language coverageeSpeak NGA few MB, 100+ languages and accents
To train and publish research modelsESPnetRecipes for many model types

Also check 3 things before you commit:

  • License of the code and of the model weights. They're often different.
  • Recent activity. Look at the date of the last release, not just the star count.
  • Your hardware. Large neural models run far faster on a GPU.

To compare output quality between toolkits, run the same script through each and score it with the listening tests in our guide to synthetic speech intelligibility metrics.

When a hosted AI voice generator makes more sense

Open-source toolkits are great for learning, privacy and custom research. They're less great when you just need a finished voiceover by this afternoon.

A hosted tool like Kveeky skips the setup: paste a script, pick a voice, download the audio. Kveeky has 700+ AI voices in 40+ languages and a free plan with 500 credits a month (about 6.6 minutes) and no credit card. Every paid plan includes commercial usage rights, which matters if XTTS's non-commercial license rules it out for you.

Where open source wins: you control the model, your audio never leaves your machine, and you can train on your own data. Where a hosted tool wins: no installs, no GPU, ready-made voices and clear commercial terms.

Frequently asked questions

Is Coqui TTS still maintained?

The original coqui-ai/TTS repo has had no release since v0.22.0 on December 12, 2023. The Idiap fork, installed as coqui-tts, is maintained. Its latest release was v0.27.5 on January 26, 2026.

Is Coqui TTS free for commercial use?

The toolkit code is MPL-2.0, which allows commercial use under its terms. Model weights have their own licenses. XTTS-v2 uses the Coqui Public Model License, which is non-commercial only, so check each model you use.

Where is the Coqui TTS GitHub repo?

The original is github.com/coqui-ai/TTS. The maintained fork is github.com/idiap/coqui-ai-TTS, and its PyPI package is coqui-tts.

Can Coqui TTS clone voices?

Yes. The XTTS-v2 model clones a voice from a short reference clip, about 6 seconds according to its model card, and supports 17 languages. Its license limits use to non-commercial projects.

What are the best open-source speech synthesis libraries?

Coqui TTS for many models and voice cloning, Piper for fast offline voices, eSpeak NG for tiny size and 100+ languages, and ESPnet for research. The best one depends on your hardware and license needs.

Do I need a GPU to run Coqui TTS?

Not always. Smaller models can run on a CPU, but large models like XTTS-v2 run much faster on a GPU. Lighter engines such as Piper or eSpeak NG are better fits for low-power machines.

How we checked this guide

This guide is written by Govind Kumar for the Kveeky team. Disclosure: Kveeky makes an AI voice generator. No Kveeky usage data is used in this guide.

Your next step: if you're new to speech synthesis, read our plain-English guide to how text-to-speech AI works before you install anything. It'll make every model name above easier to follow.

Govind Kumar
Govind Kumar

Co-Founder & CTPO

 

Govind Kumar is a product and technology leader focused on building AI-powered tools that simplify content creation for creators and marketers. His work centers on designing scalable systems that make it easier to generate, manage, and publish AI voice and audio content across modern platforms. At Kveeky, he focuses on improving product usability, automation, and AI-driven workflows that help creators produce natural-sounding voiceovers faster while maintaining quality and consistency. His approach combines technical depth with a strong emphasis on creator experience, making advanced AI capabilities accessible to everyday users. On the Kveeky blog he writes the technical guides on how text to speech works, neural TTS architectures, vocoders and voice quality.

Related Articles

Best AI Voice Generator for Small Business in 2026: 7 Paid Plans Compared
best ai voice generator for small business

Best AI Voice Generator for Small Business in 2026: 7 Paid Plans Compared

The best AI voice generator for small business in 2026, compared by real cost per finished minute, commercial rights, team seats and free-plan limits.

By Deepak Gupta October 8, 2026 15 min read
common.read_full_article
How to Start a Podcast With AI Voices: A Practical Workflow (2026)
ai podcast voice

How to Start a Podcast With AI Voices: A Practical Workflow (2026)

Use an AI podcast voice to launch your show: a 10-step checklist, script template, gear by budget, Apple and Spotify rules, and when to use your own voice.

By Mohit Singh October 7, 2026 20 min read
common.read_full_article
AI Voice for YouTube: The Complete Guide for Creators (2026)
how to use ai voice for youtube

AI Voice for YouTube: The Complete Guide for Creators (2026)

How to use AI voice for YouTube in 2026: monetization and disclosure rules from YouTube's own pages, a 7-step workflow, plus length and tone tables.

By Mohit Singh October 7, 2026 18 min read
common.read_full_article
Voice Changer vs Text to Speech vs Voice Cloning: Which One Do You Need?
voice changer vs text to speech

Voice Changer vs Text to Speech vs Voice Cloning: Which One Do You Need?

Voice changer vs text to speech vs voice cloning: what each does, latency, consent rules and a decision table for streams, calls, dubbing and voiceovers.

By Ankit Agarwal October 7, 2026 17 min read
common.read_full_article