Search papers, labs, and topics across Lattice.
This paper investigates the impact of low-frequency signals on large audio-language models (LALMs), revealing that inaudible inputs can significantly degrade model performance. The authors introduce Intermittent Low-Frequency Lockout (ILL), a novel method for evaluating these risks, and Distributional Requery Guard (DRG) to enhance model resilience against low-frequency attacks. Results show that ILL can reduce accuracy by up to 67 percentage points while maintaining low human audibility, and DRG improves attacked accuracy from 28.5% to 46.1% after clean reacquisition, highlighting a critical safety concern in LALMs.
Inaudible low-frequency signals can cripple LALM performance by up to 67%, revealing a hidden vulnerability in audio processing systems.
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting. ILL uses Sentence Attention Scale Estimation to determine active intervals and Frequency Confusion Transfer to construct a low-frequency waveform with continuous phase from corpus spectral variation. To mitigate this risk, we propose Distributional Requery Guard (DRG) to detect low-frequency distribution shifts and conditionally request a second recording for semantic recovery. Across six LALMs and multiple audio understanding tasks, ILL reduces accuracy by up to 67 percentage points while receiving a mean human audibility rating of 1.33, close to 1.17 for clean audio; DRG raises mean attacked accuracy from 28.5\% to 46.1\% after clean reacquisition. These findings identify a previously overlooked safety risk for LALMs and provide a foundation for future research on robust audio understanding.