Discover the best open-source text-to-speech, speech-to-text, voice cloning, voice activity detection and LLM models you can self-host — each with its GitHub stars, license and a link to the repo.
Submit an open-source productKyutai — Pocket TTS is a compact text-to-speech model small enough to run on a CPU — fast, lightweight, on-device voice generation.
Real-Time Voice Cloning — Clone a voice in 5 seconds to generate arbitrary speech in real-time
NVIDIA NeMo — A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
SpeechBrain — A PyTorch-based Speech Toolkit
FluidVoice — FluidVoice is an open-source voice dictation / speech-to-text tool.
Kyutai — Kyutai's open speech-to-text and text-to-speech models built on the Delayed Streams Modeling framework for streaming, low-latency speech.
FireRedTeam — FireRedASR is FireRedTeam's open-source automatic speech recognition (ASR) model for high-accuracy speech-to-text.
MiniMax — MiniMax-M3 is an open-source large language model released by MiniMax, available on GitHub to run and build on.
Fish Audio — Open-source AI voice platform with high-quality TTS and voice cloning in multiple languages.
FireRedTeam — FireRedChat is FireRedTeam's open-source real-time, full-duplex conversational voice agent.