Search papers, labs, and topics across Lattice.
This paper introduces AT-ADD, a comprehensive benchmark for evaluating robust detection of audio deepfakes across various types, including speech, environmental sounds, singing, and music. The benchmark features two tracks: one focusing on binary speech detection under diverse conditions and another on type-agnostic detection when the audio type is unknown. Key findings reveal that while the strongest baseline achieves 76.73% and 79.47% Macro-F1 scores for the respective tracks, winning challenge systems significantly outperform these with scores of 90.71% and 96.10%, highlighting the effectiveness of advanced techniques like self-supervised representations and condition-aware augmentation.
Winning systems in the AT-ADD challenge achieved over 90% accuracy in detecting all types of audio deepfakes, showcasing the potential of innovative detection strategies.
Recent audio generation models can synthesize high-fidelity speech, environmental sound, singing voice, and music, creating new risks for multimedia trust. Existing audio deepfake detection (ADD) benchmarks remain predominantly speech-centric and often underrepresent realistic channel variation and diverse audio types. This paper presents AT-ADD, a large-scale benchmark and challenge designed to evaluate both robust speech deepfake detection and all-type audio deepfake detection. Track 1 evaluates binary speech detection under unseen generators, diverse recording conditions, signal perturbations, and replay effects. Track 2 evaluates type-agnostic real/fake detection over speech, sound, singing, and music when the audio type is unknown at test time. We detail the dataset construction, evaluation protocol, and reproducible baselines, and analyze the final systems submitted to the ACM Multimedia 2026 Grand Challenge. The strongest official baseline obtains 76.73% and 79.47% Macro-F1 on the Track 1 and Track 2 evaluation sets, respectively, whereas the winning challenge systems reach 90.71% and 96.10%. Beyond aggregate rankings, sample-level analysis of the top five submissions examines generator- and type-level difficulty, cross-system error complementarity, and ranking stability. The results show that large-scale self-supervised representations, condition-aware augmentation, multi-crop inference, and structured fusion or routing are central to generalization, while generator-specific robustness and consistent performance across diverse audio types remain unresolved.