Search papers, labs, and topics across Lattice.
This paper introduces MADBench, a novel benchmark that differentiates between speech and environmental audio in the context of audio deepfake detection, addressing a significant gap in existing research that conflates these distinct acoustic components. The study reveals that while environmental audio manipulation is more detectable than synthetic speech, current state-of-the-art detectors struggle with both components, leading to a degradation in detection performance when environmental audio is manipulated. These findings underscore the inadequacy of previous benchmarks that treated audio as a single entity, highlighting the need for a more nuanced approach in deepfake detection methodologies.
Environmental audio manipulation is easier to detect than synthetic speech, but existing detectors fail to handle both effectively, revealing critical gaps in current audio forensics.
Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet existing research either focuses on visual manipulation, addresses speech detection in isolation, or conflates speech and non-speech audio as a single undifferentiated audio stream, overlooking the distinct forensic challenges posed by background audio. This conflation is consequential: the two acoustic components arise from fundamentally different generative mechanisms, exhibit distinct artifact profiles, and pose different challenges to detection systems. We introduce MADBench, the first benchmark that treats speech and environmental audio as distinct acoustic components, enabling component-aware evaluation of audio deepfake detection across independently manipulated forgery sources. We benchmark representative state-of-the-art detectors and multimodal large language models under a unified protocol. Our experiments reveal that environmental audio manipulation is more detectable than synthetic speech across general-purpose encoders, while existing pretrained detectors fail on both acoustic components, and manipulated environmental audio asymmetrically degrades speech deepfake detection, findings entirely invisible under the single-label paradigm of prior benchmarks. MADBench establishes a rigorous foundation for future research into robust, component-aware audio deepfake detection.