Search papers, labs, and topics across Lattice.
4
1
5
4
Audio-language models are blind to negation, with performance on negated sound concepts plummeting to below chance levels.
Amplifying just a handful of key neurons in the audio encoder can boost LALM accuracy on non-semantic speech attributes by over 25 points.
LALMs leak speaker identity by memorizing the link between voice and text, not just the content of speech.
Text-only LLMs already contain surprisingly diverse levels of auditory knowledge, and this pre-existing knowledge strongly predicts their performance when adapted for audio-language tasks.