Search papers, labs, and topics across Lattice.
This study reveals that audio-language embedding models, such as CLAP, exhibit a significant limitation in handling negation, as they often produce similar representations for affirmative and negated sound concepts. To investigate this issue, the authors introduce NegEval-Audio, a framework designed to evaluate models on negation-aware tasks, demonstrating that performance drops dramatically when negation is involved, with accuracy on Multiple-Choice Negation tasks falling well below chance. The findings highlight a fundamental flaw in representation geometry, suggesting that current models require explicit training on negation to improve their performance in real-world applications.
Audio-language models are blind to negation, with performance on negated sound concepts plummeting to below chance levels.
Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this affirmation-only evaluation hides a key limitation: these models fail to encode negated sound concepts, mapping affirmative and negated captions to nearly identical representations. To expose this blind spot, we introduce NegEval-Audio, a framework that converts existing datasets into two negation-aware tasks, Retrieval-Neg and Multiple-Choice Negation (MCQ-Neg), to probe whether models distinguish present from absent events. On AudioCaps and Clotho, performance degrades sharply under negation, with negation-type MCQ accuracy falling far below chance, and the failure persists even for a recent multimodal LLM-based embedding model. While a training-free steering method improves MCQ-Neg, it yields marginal gains for Retrieval-Neg. This indicates that affirmation bias is a fundamental flaw in the representation geometry, necessitating explicit negation-aware training objectives.