Search papers, labs, and topics across Lattice.
8
0
10
3
Prioritizing lexically rich narratives can accelerate ASR transcription quality improvements by overcoming the cold start problem in language documentation.
Environmental audio manipulation is easier to detect than synthetic speech, but existing detectors fail to handle both effectively, revealing critical gaps in current audio forensics.
Token-level explanations reveal how FEMRs leverage patient history, bridging the gap between black-box models and clinical trust.
LALMs struggle to remember non-speech sounds across multi-turn conversations not because of faulty attention, but because their internal representations of those sounds drift over time.
Current continual learning methods fail to account for the coupled, geometry-sensitive nature of acoustic representations in modern speech foundation models, hindering their ability to adapt to non-stationary environments.
Environmental sound deepfakes are a rising threat, and this challenge reveals the current state-of-the-art in detecting them, highlighting both the progress and remaining gaps.
LALMs struggle with polyphonic audio, losing significant performance on tasks requiring reasoning about concurrent sound events, as revealed by the new PolyBench benchmark.
LALMs get a noise-canceling superpower with Focus-Then-Listen, a plug-and-play module that boosts performance without expensive retraining.