Discover the best open-source text-to-speech, speech-to-text, voice cloning, voice activity detection and LLM models you can self-host — each with its GitHub stars, license and a link to the repo.
Submit an open-source productPatter — Open-source SDK for voice AI to connect any agent to real phone calls in 4 lines.
DeepSpeech — DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
Moonshot AI — Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
SpeechBrain — A PyTorch-based Speech Toolkit