Search papers, labs, and topics across Lattice.
This paper introduces SimulS2ST-Omni, a data-efficient framework for long-form streaming speech-to-speech translation (S2ST) that leverages only approximately 2,000 hours of paired data while employing auxiliary multitask training. The method innovatively combines joint text-code trajectory supervision to streamline the translation process and reduce reliance on separate emission controllers, which enhances stability and performance. As a result, the system achieves competitive quality-latency trade-offs, matching the performance of leading closed-source systems like LiveInterpret 2.0 on ASR-BLEU metrics.
Achieving state-of-the-art performance in streaming speech-to-speech translation with a mere 2,000 hours of paired data challenges the notion that massive datasets are essential for high-quality results.
Long-form streaming speech-to-speech translation (S2ST) requires incremental, unbounded translation under strict latency constraints. Existing methods typically suffer from sentence-bounded supervision or demand massive paired-S2ST supervision. We introduce a training recipe enabling a speech language model for sentence-level and long-form streaming S2ST using only $\sim$2k hours of paired cross-lingual S2ST data, layered atop auxiliary supervision. Anchored by auxiliary multitask training, our approach remains robust even when the paired-S2ST budget itself is reduced by 90\%. Our core contribution, joint text-code trajectory supervision, schedules target text and acoustic semantic codes as a unified commitment path, eliminating the need for separate, unstable speech-side emission controllers. Furthermore, our two-stream Thinker--Talker factorization significantly outperforms unified-decoder baselines by decoupling linguistic reasoning from dense acoustic prediction to mitigate modality interference. Finally, our system achieves highly competitive quality-latency trade-offs on RealSI and ACL60/60-dev, matching state-of-the-art, closed-source S2ST systems such as LiveInterpret~2.0 on ASR-BLEU.