Search papers, labs, and topics across Lattice.
This paper introduces AffectDF, a new benchmark for speech deepfake detection that comprehensively addresses the challenges posed by emotionally expressive attacks, including various spoofing methods such as text-to-speech and voice conversion. The study reveals that existing state-of-the-art detection systems exhibit significant performance degradation when evaluated against AffectDF, often nearing random guessing, particularly under emotional variability. Notably, even extensive emotional training fails to enhance cross-domain robustness, highlighting critical gaps in current detection methodologies and the need for improved models that can generalize across diverse spoofing scenarios.
Current speech deepfake detection systems falter dramatically against emotionally expressive attacks, with performance dropping to near-random levels on the new AffectDF benchmark.
Speech deepfake detection (SDD) systems achieve strong performance on conventional benchmarks; however, existing datasets provide limited coverage of emotionally expressive and recent large audio-language model (LALM)-based attacks. Existing emotional spoofing datasets are also limited in scale and attack diversity, typically covering only voice conversion (VC) or text-to-speech (TTS) attacks. We introduce AffectDF, the most comprehensive benchmark for emotionally expressive speech deepfakes, spanning TTS, VC, emotional VC, and LALM-based spoofing attacks across both acted and spontaneous emotional speech. AffectDF contains approximately 260 hours of speech generated using 21 spoofing attacks across five emotional states. We benchmark state-of-the-art SDD systems under conventional and emotional spoofing conditions, including LALM-based detectors evaluated with both inference-only prompting and supervised fine-tuning. Our experiments reveal severe robustness degradation when models trained on conventional benchmarks are evaluated on AffectDF, with several systems approaching near-random performance. Surprisingly, even large-scale emotional training does not consistently improve cross-domain robustness, indicating that current SDD systems fail to learn generalized spoof representations under emotional and prosodic variability. Robustness further varies substantially across emotional states, attack families, and acted vs spontaneous emotional speech conditions. These findings expose fundamental limitations of current SDD systems and establish AffectDF as a benchmark for developing more robust spoof detection models.