Search papers, labs, and topics across Lattice.
This paper introduces Evidence Subspace Projection, a novel method that quantitatively connects the features captured by self-supervised learning (SSL) models to their decisions in audio deepfake detection. By projecting decision vectors onto a shared space of evidence factors, the authors derive a scalar ratio that measures the explanatory power of various evidence types, such as attack category and codec. The evaluation across multiple datasets not only validates the method against established findings but also uncovers new insights into the behavior of SSL models in different settings.
Evidence Subspace Projection reveals how different factors influence deepfake detection decisions, providing a clearer understanding of model behavior in self-supervised learning.
Self-supervised learning (SSL) models are widely used as feature extractors for state-of-the-art audio deepfake detection, but it remains unclear how to directly and quantitatively connect what SSL models capture to detection decisions. To address this gap, we propose Evidence Subspace Projection, a method that represents both evidence factors (e.g., attack category, codec, gender, transmission) and authenticity labels in a shared space constructed from SSL models'neuron activation patterns. By projecting the decision vector onto each evidence subspace, we obtain a scalar ratio that quantifies the explanatory power of each evidence type. We evaluate SSL models in raw, fine-tuned, and post-trained settings on multiple datasets. The results confirm findings from established studies, validating the proposed method, and reveal new insights into model behavior.