Best Open Source tts Libraries
A curated list of the most popular GitHub repositories tagged with tts. Select any project to visualize its architecture and dive into the codebase using RepoMind's AI engine.
#1CorentinJ/Real-Time-Voice-Cloning
Clone a voice in 5 seconds to generate arbitrary speech in real-time
#2RVC-Boss/GPT-SoVITS
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
#3unslothai/unsloth
Fine-tuning & Reinforcement Learning for LLMs. 🦥 Train OpenAI gpt-oss, DeepSeek, Qwen, Llama, Gemma, TTS 2x faster with 70% less VRAM.
#4coqui-ai/TTS
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
#5mudler/LocalAI
:robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more. Features: Generate Text, MCP, Audio, Video, Images, Voice Cloning, Distributed, P2P and decentralized inference
#62noise/ChatTTS
A generative speech model for daily dialogue.
#7babysor/MockingBird
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
#8myshell-ai/OpenVoice
Instant voice cloning by MIT and MyShell. Audio foundation model.
#9fishaudio/fish-speech
SOTA Open Source TTS
#10mastra-ai/mastra
Mastra is the modern TypeScript framework for AI-powered applications and agents.
#11index-tts/index-tts
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
#12readest/readest
Readest is a modern, feature-rich ebook reader designed for avid readers offering seamless cross-platform access, powerful tools, and an intuitive interface to elevate your reading experience.
#13DrewThomasson/ebook2audiobook
Generate audiobooks from e-books, voice cloning & 1158+ languages!
#14pot-app/pot-desktop
🌈一个跨平台的划词翻译和OCR软件 | A cross-platform software for text translation and recognition.
#15NVIDIA-NeMo/NeMo
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
#16moonshine-ai/moonshine
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
#17krillinai/KrillinAI
Video translation and dubbing tool powered by LLMs. The video translator offers 100 language translations and one-click full-process deployment. The video translation output is optimized for platforms like YouTube,TikTok. AI视频翻译配音工具,100种语言双向翻译,一键部署全流程,可以生抖音,小红书,哔哩哔哩,视频号,TikTok,Youtube等形态的内容成适配
#18fishaudio/Bert-VITS2
vits2 backbone with multilingual-bert
#19dramaclaw/dramaclaw
A general-purpose AIGC video engine: script to finished film in one pipeline — dramas, ads, product videos, otome games, and more. | 通用 AIGC 视频引擎 —— 从剧本到成片一条流水线,漫剧、广告、电商、乙游皆可
#20kizuna-ai-lab/sokuji
Real-time two-way speech translation for bilingual meetings — auto-detects the spoken language and translates both directions, cloud or fully offline on-device. Desktop (Windows · macOS · Linux) + browser extension (Chrome · Edge) for Zoom, Meet, Teams & any app.
#21sgl-project/sglang-omni
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
#22ChenShuo2004/cs-board
将参考声音和中文文案自动生成白板动画视频的本地 AI 工具。