Search papers, labs, and topics across Lattice.
This paper introduces an interpretable framework for detecting speech deepfakes using artifact-specific expert models that provide human-understandable evidence rather than relying on opaque decision-making. Each expert is trained to identify a specific artifact associated with speech synthesis, and their outputs are calibrated into log-likelihood ratios that serve as interpretable evidence scores. The evaluation demonstrates that these experts effectively capture their respective artifacts and contribute to a robust ensemble classification, enhancing interpretability in high-stakes scenarios.
Artifact-specific experts can provide interpretable evidence in speech deepfake detection, transforming how we assess synthetic audio authenticity.
In this work, we propose an interpretable framework for speech deepfake detection based on artifact-specific expert models. Rather than relying on black-box decisions, the framework provides human-understandable evidence, which is critical in high-stakes settings. Each expert is trained to detect a specific speech synthesis artifact, and its output is calibrated into a log-likelihood ratio that serves as an interpretable evidence score. We evaluate five artifact-specific experts and show that, with proper calibration, they can capture their target artifacts and produce meaningful evidence. Importantly, each expert estimates only the presence of its assigned artifact rather than directly performing the final decision. Their outputs are aggregated into an ensemble to produce the actual real-versus-fake classification, while maintaining interpretability by indicating how strongly each expert supports or contradicts a fake classification. Results show that artifact-specific experts capture interpretable signals of synthetic speech across multiple generation pipelines.