Search papers, labs, and topics across Lattice.
This study explores the embedding of a 32-bit watermark into the continuous latent representation of a neural audio codec, specifically a speech autoencoder, to enhance codec robustness against transformations. By employing a SEANet-style encoder-decoder and a Conformer-based message embedder, the authors analyze the trade-offs associated with watermarking in the latent domain, rather than traditional waveform or spectrogram methods. The results show significant improvements in bit accuracy, with EnCodec-aware training boosting accuracy from 78.8% to 95.6% and 97.1%, albeit with a slight decrease in PESQ scores.
Embedding watermarks in the latent space of neural codecs can drastically improve robustness, achieving up to 97.1% accuracy in challenging conditions.
Neural audio codecs are challenging transformations for audio watermarking because they re-encode, quantize, and resynthesize speech. This paper investigates continuous latent-space watermarking for codec robustness. Instead of adding a watermark only to the waveform or spectrogram, we embed a 32-bit message into the continuous latent representation of a codec-like speech autoencoder. The pipeline uses a SEANet-style encoder-decoder, a Conformer-based message embedder, RVQ-guided latent decomposition, and a latent-domain detector trained under signal-processing and neural-codec transformations. Rather than proposing a final universal watermarking baseline, we characterize the trade-offs that appear when the watermark carrier is moved before neural decoding. On 48 kHz speech, EnCodec-aware training improves EnCodec-24k bit accuracy from 78.8% to 95.6% and 97.1%, while PESQ decreases from 3.727 to 3.514 and 3.427.