
Kyutai
Open-science AI lab in Paris building open speech models — Moshi, TTS, STT, and translation.

About Kyutai
Kyutai is an open-science AI research lab based in Paris, on a mission to build and democratize AI through open science. It develops and openly releases state-of-the-art speech and voice models — including the Moshi full-duplex spoken-dialogue foundation model, streaming speech translation, and compact TTS/STT — freely available on GitHub and Hugging Face.
Open-source projects
- Moshi — speech-text foundation model and full-duplex spoken-dialogue framework (with the Mimi neural audio codec)
- Pocket TTS — a text-to-speech model small enough to run on a CPU
- Delayed Streams Modeling — Kyutai's streaming speech-to-text and text-to-speech models
- Hibiki — real-time, streaming speech translation
Who Is It For?
- Researchers and developers building real-time voice AI
- Teams wanting open, self-hostable speech models
- The open-source and open-science community
Use Cases
- Real-time, full-duplex voice conversation (Moshi)
- Streaming STT and TTS, including on-device TTS
- Live speech-to-speech translation (Hibiki)
Explore Kyutai's open speech models on GitHub and Hugging Face.