Search papers, labs, and topics across Lattice.
This study introduces the Perceptible-Imperceptible Passive-fingerprint Diagnostic Protocol (PIPDP) to analyze and differentiate between perceptible and imperceptible fingerprints in speech deepfake detection. The findings reveal that while perceptible fingerprints can be influenced by emotional expression and model updates, imperceptible fingerprints offer more consistent and reliable attribution cues across various speech generators and detectors. Notably, perceptually transparent perturbations significantly decrease attribution accuracy, underscoring the robustness of imperceptible fingerprints in maintaining attribution reliability.
Imperceptible fingerprints prove to be far more reliable for speech deepfake attribution, even as perceptible ones fluctuate with emotional content and model changes.
Passive fingerprints (intrinsic traces naturally left by generators) have been shown to enable attribution in speech deepfake detection, yet their persistence, reproducibility, and content-independence remain unverified. Moreover, no prior work distinguishes perceptible from imperceptible fingerprints, although the two have very different implications for attribution reliability. Perceptible fingerprints, such as emotional expression, are shaped by perceptual quality objectives and may change across model updates, whereas imperceptible fingerprints are not explicitly optimised by current training objectives and are rarely considered in existing dataset design or training strategies, as they have limited influence on downstream applications. We therefore propose a Perceptible-Imperceptible Passive-fingerprint Diagnostic Protocol (PIPDP) to define and separately analyze these two fingerprint types. PIPDP comprises three complementary analyses: multi-evidence fingerprint verification through residual-energy, reproducibility, and saliency analyses, perceptually transparent perturbations preserving audio quality, and prompt-driven emotion change that modifies perceptible fingerprints without model retraining. Experiments across ten speech generators and three attribution detectors show that imperceptible fingerprints provide persistent attribution cues. Perceptually transparent perturbations reduce attribution accuracy by up to 48.2\% on HiggsAudioV3, whereas emotion-driven changes leave attribution largely unchanged, with only about a 1.0\% accuracy variation across emotions on CosyVoice2 using w2v-bert-MLP. These results suggest that imperceptible fingerprints are more reliable for trustworthy attribution.