Discover the best open-source text-to-speech, speech-to-text, voice cloning, voice activity detection and LLM models you can self-host — each with its GitHub stars, license and a link to the repo.
Submit an open-source productMoonshot AI — Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
Speechmatics — Open-source AI speech tech for enterprise, offering real-time transcription, translation, and TTS.
NVIDIA NeMo — A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
Fish Audio — Open-source AI voice platform with high-quality TTS and voice cloning in multiple languages.
Kyutai — Pocket TTS is a compact text-to-speech model small enough to run on a CPU — fast, lightweight, on-device voice generation.
KugelAudio — Europe's first production-ready TTS with 40+ languages, developed and hosted in Europe, fully GDPR compliant.
FluidVoice — FluidVoice is an open-source voice dictation / speech-to-text tool.
Real-Time Voice Cloning — Clone a voice in 5 seconds to generate arbitrary speech in real-time
Silero VAD — Silero VAD: pre-trained enterprise-grade Voice Activity Detector
pyannote.audio — Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding
LobbyStack — Open-source AI receptionist that answers calls, qualifies leads, books appointments, and routes urgent requests 24/7.
FireRedTeam — FireRedTTS2 is FireRedTeam's open-source text-to-speech model for natural, expressive voice generation.