Search papers, labs, and topics across Lattice.
This paper identifies a novel vulnerability in invisible watermarking through the lens of foundation image models, termed "watermark laundering," where a single reconstruction prompt can effectively erase watermarks from images. The authors evaluate this phenomenon using a joint payload-fidelity profile, revealing that OpenAI models exhibit significant payload disruption while other models like Nano Banana 2 remain susceptible under high-fidelity conditions. Their findings highlight that the reconstruction pathway, rather than specific prompt wording, is the primary driver of this vulnerability, underscoring the need for enhanced robustness evaluations in watermarking techniques.
A single prompt can completely strip invisible watermarks from images, revealing a critical vulnerability in current watermarking schemes.
Invisible watermarks are typically evaluated against predefined perturbations such as compression, blur, noise, cropping, and denoising. Public foundation image models expose a distinct threat: an attacker can submit a watermarked image with a single reconstruction prompt and obtain a visually faithful output from which the invisible watermark can no longer be decoded reliably. We formalize this failure mode as watermark laundering and evaluate it using a joint payload-fidelity profile that combines bit error rate (BER) with visual and semantic preservation. Across six OpenAI and Google image editing models, three representative watermarking schemes, and 1,800 reconstructed outputs, we identify two complementary laundering regimes: OpenAI models produce the strongest payload disruption across the evaluated schemes, whereas Nano Banana 2 shows that DwtDct remains vulnerable under high-fidelity reconstruction. Prompt ablations show that no single removal-oriented instruction is necessary for payload disruption, indicating that the effect is primarily induced by the reconstruction pathway rather than by explicit attack wording. Comparisons with conventional attacks further show that prompt-conditioned reconstruction constitutes a distinct operational attack interface. These findings motivate foundation-model reconstruction as a missing robustness condition in invisible watermark evaluation.