Search papers, labs, and topics across Lattice.
This paper re-examines continual learning (CL) in speech and audio, arguing that the field needs to shift its focus from task-specific knowledge retention to the evolution of shared representation geometry in modern speech foundation models. They introduce a representation-centric taxonomy for CL in speech, categorizing approaches based on how they manage the geometry of entangled acoustic representations under non-stationary conditions. The authors then highlight mismatches between current CL assumptions and speech foundation model behavior, outlining open challenges for future research.
Current continual learning methods fail to account for the coupled, geometry-sensitive nature of acoustic representations in modern speech foundation models, hindering their ability to adapt to non-stationary environments.
Speech and audio systems operate in inherently non-stationary environments, yet continual learning (CL) research in this domain, especially in the foundation model era, remains fragmented that fail to account for the coupled, geometry-sensitive nature of acoustic representations. Modern speech foundation models operate over highly entangled, continuous representations that jointly encode linguistic, speaker, and paralinguistic factors within a shared latent space. CL is therefore fundamentally about preserving and evolving shared representation structure rather than retaining isolated task knowledge. In this work, we revisit CL for speech from a representation-centered perspective, and introduce a new taxonomy that organizes CL according to how underlying representation geometry evolves under non-stationary acoustic conditions. We further identify key mismatches between current CL assumptions and speech foundation model behavior, and finally outline a set of open challenges and future research directions.