Search papers, labs, and topics across Lattice.
This paper introduces a Mixture-of-Experts (MoE) framework for Automatic Speech Recognition (ASR) that effectively addresses the challenges of recognizing speech from both children and adults across diverse environments. By employing a Classifier-based Domain Router (C-DR) with a coarse-to-fine strategy and integrating Mixture-of-Projectors (MoP) and Mixture-of-LoRAs (MoL), the model captures domain-specific variations while an Entropy-Aware Routing (EAR) mechanism mitigates routing uncertainty. Experimental results on public child corpora show significant performance improvements over baseline models, indicating the framework's capability to unify ASR for different age groups without compromising adult performance.
Unified ASR for both children and adults is now achievable with a novel Entropy-Aware routing mechanism that enhances performance across diverse speech domains.
While Speech Large Language Models (Speech-LLMs) have achieved strong performance on adult Automatic Speech Recognition (ASR), their effectiveness on child speech remains under-explored, and single models often struggle to handle diverse adult and child age groups simultaneously. This paper proposes a Mixture-of-Experts (MoE) Speech-LLM for unified ASR across adult and child speech spanning diverse environments and age groups. The framework employs a Classifier-based Domain Router (C-DR) with a coarse-to-fine strategy and integrates both a Mixture-of-Projectors (MoP) and a Mixture-of-LoRAs (MoL) to model domain-specific variations. To address routing uncertainty near domain boundaries, an Entropy-Aware Routing (EAR) mechanism is introduced to dynamically incorporate a shared expert. Experiments on public child corpora demonstrate consistent improvements over baselines while preserving adult ASR performance. To our knowledge, this is the first work leveraging Speech-LLMs for unified, multi-domain ASR encompassing both children and adults.