Kyutai Releases Hibiki-Zero: A3B Parameter for Simultaneous Speech-to-Speech Translation Using GRPO Reinforcement Learning Without Any Word-Level Aligned Data
Kyutai released Hibiki-Zeroa new model for simultaneous speech-to-speech translation (S2ST) and speech-to-text translation (S2TT). The system translates the source speech into the target language in real time. It handles non-monotonic term dependencies in the process. Unlike previous models, Hibiki-Zero does not require word-level aligned data for training. This eliminates a major bottleneck in scaling AI … Read more