Search papers, labs, and topics across Lattice.
This paper introduces a dataset-aware framework for audio deepfake detection that leverages multitask learning and gradient reversal loss to enhance generalization across diverse datasets. By utilizing dataset identity as a supervisory signal, the approach circumvents the limitations of traditional methods that rely on auxiliary annotations, achieving significant performance improvements. Experimental results show a 13.14% relative reduction in Average Equal Error Rate (EER) and a 5.32% reduction in Pooled EER compared to baseline models, indicating its effectiveness in real-world applications.
Dataset-aware multitask learning can slash deepfake detection errors by over 13%, even without auxiliary annotations.
Recent advances in speech synthesis and voice conversion, which pose threats to security and privacy, have underscored the need for deepfake detection technology. Although existing detection systems achieve strong performance on individual datasets, they often fail to generalize across diverse datasets. Prior methods for improving generalization, including data augmentation, adversarial training on auxiliary factors such as language or codec types, and Mixture-of-Experts (MoE), are limited by predefined augmentation coverage, difficulties in obtaining auxiliary factors, and substantial model complexity. In this work, we propose a practical dataset-aware framework for deepfake detection. Our method targets heterogeneous datasets for which auxiliary annotations such as language, codec, or spoofing method may not be consistently available. We therefore rely only on dataset identity as a naturally available supervisory signal for multitask (MT) and gradient reversal layer (GRL) training, allowing the model to investigate both dataset-aware multitask supervision and adversarial suppression of dataset-specific information. We conduct experiments following the 2025 Speech DeepFake Arena benchmark protocol, evaluating our model across multiple evaluation datasets and reporting aggregate performance in terms of Equal Error Rate (EER), including Average EER and Pooled EER. Compared with the baseline, MT reduces Average EER by 13.14% relatively, while GRL reduces Pooled EER by 5.32% relatively. These results demonstrate that our method can improve aggregate detection performance across heterogeneous evaluation datasets, offering a practical solution for deploying reliable deepfake detection systems on diverse and unseen real-world data.