Search papers, labs, and topics across Lattice.
This paper introduces VoxSumm, a pioneering multilingual corpus designed for joint speech summarization and translation (JSumT), addressing the gap between long-document summarization and multilingual speech processing. By compiling 10,045 BBC article-summary pairs across 24 languages and approximately 703 hours of speech data, the authors provide a robust benchmark for evaluating models in this domain. Their findings indicate significant variability in performance across different models and settings, with Gemini3.1-Pro showing the highest consistency, and highlight challenges in translating before summarization that lead to instruction-following failures.
Jointly summarizing and translating long-form spoken content could revolutionize how we handle multilingual information processing.
As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, long-document summarization research remains text-centric, whereas multilingual speech research has largely prioritized translation, preserving source content rather than compressing it. We address this methodological gap by formalizing joint speech summarization and translation (JSumT): the generation of a succinct, faithful target-language summary directly from a long spoken document in a source language. We additionally introduce VoxSumm, the first multilingual and cross-lingual benchmark for this task, comprising 10,045 BBC article-summary pairs across 24 languages and encompassing approximately 703 hours of speech data. Our evaluation of representative speech-language models reveals pronounced variation across models and generation settings: Gemini3.1-Pro demonstrates the greatest consistency, summarization into English generally surpasses generation into non-English target languages, and translating an entire document before summarization compounds instruction-following failures. Through the release of VoxSumm, we establish a foundation for developing and evaluating multilingual systems capable of jointly interpreting, compressing, and translating long-form speech.