Search papers, labs, and topics across Lattice.
This paper introduces Reference-Augmented Training (RAT), a novel architecture for anti-spoofing in automatic speaker verification (ASV) that utilizes speaker-reference recordings. The key finding reveals that while the model initially relies on the reference during training, it ultimately learns to perform effectively without it, achieving significant improvements in deepfake detection. RAT achieves state-of-the-art performance on the ASVspoof 5 benchmark, with a 2.57% equal error rate (EER) and 0.074 minimum detection cost function (minDCF), outperforming even large ensemble systems.
Training with speaker references might seem essential, but RAT shows that models can excel in deepfake detection even when those references are absent during inference.
We introduce a spoofing countermeasure architecture conditioned on speaker-reference recordings, but observe that it converges to a solution that effectively ignores the reference during inference. Surprisingly, training with a reference channel induces invariance that improves deepfake detection, even when the reference is absent or mismatched during inference. Based on this observation, we propose a Reference-Augmented Training (RAT) strategy. RAT yields improved detection performance compared to single-utterance baselines, even when the reference recording is replaced with a zero vector at inference. Through rigorous analysis, we demonstrate that the optimization process rapidly diminishes the reference contributions, leading to inference largely independent of the reference channel. Using RAT, we achieve state-of-the-art 2.57% EER and 0.074 minDCF on the ASVspoof 5 benchmark with a single detector, surpassing even large ensemble systems.