Discover the best open-source text-to-speech, speech-to-text, voice cloning, voice activity detection and LLM models you can self-host — each with its GitHub stars, license and a link to the repo.
Submit an open-source productDeepSpeech — DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
Common Voice — Mozilla's open, crowdsourced multilingual speech dataset for training speech recognition — millions of validated voice clips with transcripts across 100+ languages, released under CC0. The platform code is MPL-2.0.
FluidVoice — FluidVoice is an open-source voice dictation / speech-to-text tool.
FireRedTeam — FireRedChat is FireRedTeam's open-source real-time, full-duplex conversational voice agent.
LobbyStack — Open-source AI receptionist that answers calls, qualifies leads, books appointments, and routes urgent requests 24/7.
Kyutai — Pocket TTS is a compact text-to-speech model small enough to run on a CPU — fast, lightweight, on-device voice generation.
KugelAudio — Europe's first production-ready TTS with 40+ languages, developed and hosted in Europe, fully GDPR compliant.
SpeechBrain — A PyTorch-based Speech Toolkit