Browse, search and filter Voice AI companies.
minimax.io
Open-source multi-modal models powering AGI with breakthrough capabilities.
soniox.com
Multilingual Speech AI API for real-time transcription, translation, and speech synthesis with sub-200ms latency.
deepl.com
Real-time, secure voice translation for global teams to communicate effortlessly across languages.
inworld.ai
High-quality, real-time Voice AI with voice cloning and emotional expressiveness.
krisp.ai
Voice AI for meetings with noise cancellation, AI notes, accent conversion, and call center AI.
ai-coustics.com
Real-time audio enhancement for reliable voice AI in production environments.
vonage.com
VoIP and Unified Communications Solutions for Business Connectivity
resemble.ai
Secure your media with multimodal deepfake detection and watermarking from Resemble AI.
theten.ai
Open-source real-time voice-agent framework — home of TEN VAD, a fast, lightweight voice activity detector.
github.com/pyannote
Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech d
github.com/snakers4
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
github.com/corentinj
Clone a voice in 5 seconds to generate arbitrary speech in real-time
github.com/meta-llama
Inference code for Llama models
github.com/google-deepmind
Gemma open-weight LLM library, from Google DeepMind
github.com/deepseek-ai
DeepSeek-V3 — open source.
github.com/zai-org
GLM-5: From Vibe Coding to Agentic Engineering
github.com/julius-speech
Open-Source Large Vocabulary Continuous Speech Recognition Engine
github.com/kaldi-asr
kaldi-asr/kaldi is the official location of the Kaldi project.
github.com/espnet
End-to-End Speech Processing Toolkit
github.com/mozilla
DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices r
github.com/alphacep
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
github.com/speechbrain
A PyTorch-based Speech Toolkit
github.com/moonshine-ai
Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces
github.com/openai
Robust Speech Recognition via Large-Scale Weak Supervision
github.com/qwenlm
Qwen3-ASR is an open-source series of ASR models developed by the Qwen team at Alibaba Cloud, supporting stable multilin
github.com/nvidia-nemo
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, an
github.com/zyphra
Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech,
github.com/funaudiollm
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
github.com/swivid
Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"
github.com/2noise
A generative speech model for daily dialogue.