Search papers, labs, and topics across Lattice.
The University of Melbourne
5
0
8
3
LALMs struggle to remember non-speech sounds across multi-turn conversations not because of faulty attention, but because their internal representations of those sounds drift over time.
Current continual learning methods fail to account for the coupled, geometry-sensitive nature of acoustic representations in modern speech foundation models, hindering their ability to adapt to non-stationary environments.
LALMs get a noise-canceling superpower with Focus-Then-Listen, a plug-and-play module that boosts performance without expensive retraining.
Environmental sound deepfakes are a rising threat, and this challenge reveals the current state-of-the-art in detecting them, highlighting both the progress and remaining gaps.
LALMs struggle with polyphonic audio, losing significant performance on tasks requiring reasoning about concurrent sound events, as revealed by the new PolyBench benchmark.