Search papers, labs, and topics across Lattice.
This paper systematically investigates the application of federated learning (FL) to SpeechLLM-based end-to-end automatic speech recognition (ASR) systems, addressing the challenges of high-dimensional parameter spaces and communication overhead. By implementing a tailored communication-efficient federated optimization strategy, the authors demonstrate that their approach achieves competitive word error rates while significantly reducing communication costs in both English and Italian ASR tasks. The findings establish a practical foundation for deploying federated SpeechLLMs in real-world multilingual environments, highlighting the potential for privacy-preserving ASR systems.
Federated learning can enhance SpeechLLM performance while slashing communication costs, paving the way for privacy-preserving ASR in diverse settings.
Federated learning (FL) enables privacy-preserving training of automatic speech recognition (ASR) systems across distributed data sources, yet its application to large-scale speech language models (SpeechLLMs) remains unexplored. This paper presents the first systematic study of federated training for SpeechLLM-based end-to-end ASR systems. We design a communication-efficient federated optimization strategy tailored to the unique challenges of SpeechLLM architectures, addressing high-dimensional parameter spaces, gradient communication overhead, and computational constraints in distributed settings. Through extensive empirical evaluation on monolingual ASR tasks in English and Italian, we demonstrate the effectiveness and stability of our federated approach compared to centralized training baselines across diverse acoustic conditions and speaking styles. Additionally, we conduct a comprehensive ablation study analyzing the impact of different speech encoder architectures on monolingual English ASR performance within the federated framework, providing insights into optimal model configurations for decentralized training. Our results achieve competitive word error rates while reducing communication costs, establishing practical foundations for federated SpeechLLM deployment in real-world multilingual scenarios.