Search papers, labs, and topics across Lattice.
This paper introduces AudioNoisePrints, a novel watermarking technique for flow matching and diffusion text-to-speech (TTS) models that operates without requiring model retraining or compromising output quality. By leveraging the inherent spatial correlations between initial Gaussian noise and generated audio, the authors demonstrate that a simple cosine correlation can effectively embed watermarks. The proposed method outperforms the existing AudioSeal baseline, indicating its robustness across various TTS and vocoder models, and suggests broader applicability for future audio generation technologies.
Watermarking TTS outputs without retraining or quality loss is now feasible, thanks to a novel method that exploits noise correlations.
We present AudioNoisePrints, a training-free watermarking pipeline for flow matching and diffusion TTS models, which requires minimal extra computation during inference and does not require retraining the TTS model or reducing the generation quality. We exploited the fact that there are strong correlations between the initial Gaussian noises and the generated outputs in diffusion and flow matching models, such that a simple cosine correlation between the initial noise and the generated output can be used to perform watermaking. Moreover, we train a lightweight detector on top for more aggressive augmentations. Our method outperforms AudioSeal, a strong baseline for audio watermarking under strong augmentations. We experimented on F5TTS and other TTS and vocoder models, and concluded that they all exhibit similar spatial correlation properties, suggesting our watermarking scheme can be used for more flow-matching TTS models and even vocoders in the future.