Search papers, labs, and topics across Lattice.
This paper introduces a novel approach to combat partial deepfake speech manipulation through self-embedding steganography, which allows clean audio to embed a compressed representation of itself for later reference extraction. The method addresses the limitations of existing passive detection systems, which struggle as the proportion of manipulated audio decreases. Experimental results demonstrate that this training-free technique effectively enhances detection capabilities and complements traditional defenses, showcasing its robustness and data efficiency.
Self-embedding steganography can transform clean speech into a resilient defense against partial deepfake manipulation without requiring any training.
Partial deepfake speech, where only limited segments of an utterance are synthesized or manipulated, poses a significant challenge to existing deepfake detection systems. As the proportion of spoofed regions decreases, passive detectors become increasingly unreliable, and accurate detection and restoration remain challenging. In this paper, we revisit audio steganography from a new perspective and propose its use as a proactive defense against partially deepfaked audio. In particular, we consider a self-embedding strategy in which a clean speech signal embeds a compressed representation of itself, enabling post-hoc extraction of reference content. We demonstrate how existing audio steganography methods can be repurposed to support detection of partial deepfakes through codec-based restoration. Experiments on a benchmark dataset show that the proposed approach complements passive defenses. Remarkably, the proposed method operates without any training, providing a robust and data-efficient alternative for partial deepfake detection.