Search papers, labs, and topics across Lattice.
This paper explores a novel approach to audio watermarking by combining self-embedding techniques with ultra-low-bitrate neural codecs to enhance content integrity verification in manipulated speech recordings. The authors demonstrate that their method not only allows for reliable detection and localization of edits but also enables the recovery of manipulated segments, overcoming limitations of traditional hash-based schemes. Experimental results confirm that the embedded payload is consistently recovered without errors, highlighting the critical role of the chosen neural codec in performance outcomes.
Embedding a neural codec representation allows for the recovery of manipulated audio segments, a breakthrough for content integrity in speech recordings.
Partial manipulation of speech recordings, where only localized segments of an utterance are altered, poses a significant challenge for content integrity verification, as reliable detection and localization of such edits becomes harder as the manipulated proportion decreases. Watermarking offers a proactive defense alternative by embedding auxiliary information prior to distribution; classical hash-based schemes achieve near-perfect detection and localization under ideal conditions, but the original content cannot be recovered once a segment is manipulated. Building on a prior self-embedding audio steganography framework, this work presents an initial exploration of proactive defense performance under ideal conditions, extending the investigation along three axes: frame-level localization, multi-bit least significant bit variants, and evaluation across multiple ultra-low-bitrate neural codec representations. By embedding a compact neural codec representation rather than a cryptographic hash, the framework additionally enables recovery of the manipulated regions, while supporting training-free detection and localization without spoofed examples. Experiments across four controlled manipulation types under ideal channel conditions show that the embedded payload, and hence an approximate reconstruction of the authentic content, is always fully recovered without bit errors. The results also indicate that the choice of neural codec is the dominant factor for detection and localization performance.