NAVER LABS Europe seminars are open to the public. Please register to participate.
Date: 16th September 2026, 11:00 am (CEST)
Abstract: In 2024, Kyutai introduced Moshi, the first full-duplex speech-to-speech conversational model capable of listening, speaking, and writing simultaneously. This release marked the beginning of a series of models built around the technical foundations of what is now known as the Moshi framework, including Hibiki, Moshi, MoshiVis, Kyutai-ASR/Kyutai-TTS, MoshiRAG and Hibiki-Zero. These models address a wide range of audio and multimodal tasks, while sharing a common requirement: real-time audio understanding and generation. By combining several technical innovations including Residual Vector Quantization (RVQ), a dual-Transformer architecture, delayed-stream modeling, and post-training techniques, the Moshi framework enables models to process audio, text, and/or vision inputs and handle multiple sequences in parallel (batch inference), while remaining below 7B parameters. This makes the approach suitable for consumer-grade hardware and even on-device applications. The presentation will first introduce the Moshi framework and its key technical components, then take a deeper dive into speech translation through Hibiki and Hibiki-Zero. It will conclude with a brief overview of MoshiVis and MoshiRAG, highlighting the broader capabilities and applications of the Moshi framework.

