Search papers, labs, and topics across Lattice.
This paper details a three-year collaboration with Phonexia to address the challenges faced in deploying deepfake speech detection systems, highlighting the significant performance drop when confronted with unseen attacks and distribution shifts. The authors emphasize that existing public benchmarks are often unsuitable for commercial applications, as they do not reflect the complexities of real-world inputs, which are typically longer and more degraded than standard test clips. Instead of introducing a new detection model, the paper advocates for the establishment of shared standards for datasets, realistic benchmarks, and interpretable scoring systems to enhance the practical deployment of detection technologies.
Deepfake speech detection may achieve sub-1% error rates in controlled settings, but real-world performance falters dramatically due to unforeseen challenges.
Synthetic speech detection benchmarks now report sub-1% error rates on some in-domain evaluations, yet performance degrades under unseen attacks, channel mismatch, and distribution shift. Based on a three-year effort with Phonexia, a commercial speaker-recognition vendor, we report barriers encountered while building and deploying a detector. Many public benchmarks are not licensed for commercial model development. Real inputs are not four-second clean clips but long, codec-degraded, sometimes partially synthetic recordings. And when a calibrated system returns a log-likelihood ratio of 2.5, no one can tell the customer what it means for their decision. Rather than proposing a new model, we connect these barriers to concrete research and coordination proposals: shared standards for commercially usable datasets, realistic deployment benchmarks, and scores that non-experts can act on. These observations come from one project and should be tested in other settings.