Search papers, labs, and topics across Lattice.
This paper introduces a fluency-aware optimization framework for simultaneous speech-to-speech translation that addresses the disruptive pauses often caused by low-latency translation methods. By leveraging model-internal signals such as linguistic diversity and temporal variability, the framework minimizes inter-chunk silences, resulting in a more natural acoustic flow. Experiments demonstrate that this approach achieves a balance between maintaining competitive latency and enhancing the naturalness of speech output, significantly reducing cognitive load for listeners.
NaturalFlow reduces disruptive pauses in simultaneous translation, achieving smoother speech flow without sacrificing latency or translation quality.
Simultaneous speech-to-speech translation aims to enable near-real-time communication by minimizing latency, offering a compelling, real-time alternative to the high latency of consecutive translation. However, the excessive pursuit of low latency often results in fragmented chunk-wise speech. Consequently, listeners are subjected to an unnatural acoustic flow punctuated by frequent pauses, which could increase their cognitive load. To bridge this gap, we introduce a fluency-aware optimization framework designed to discover the sweet spot between the low-latency benefits of simultaneous translation and the natural flow of consecutive translation. Our framework minimizes inter-chunk silences by leveraging model-internal signals, including linguistic diversity and induced temporal variability in speech durations. Experiments on short- and long-form benchmarks show that our framework produces natural speech flow while maintaining competitive latency and translation quality.