Search papers, labs, and topics across Lattice.
This study investigates the adversarial robustness of reconstruction-based AI-generated image detectors, revealing their susceptibility to crafted adversarial examples that exploit autoencoder reconstruction errors. By developing two novel attack methods, the authors demonstrate that imperceptible adversarial perturbations can significantly degrade detection performance, misclassifying synthetic images as real. The findings highlight a critical vulnerability inherent to these detectors, as adversarial examples transfer effectively across different models, raising concerns about their reliability in practical applications.
Adversarial examples can exploit reconstruction-based detectors, leading to a significant drop in detection accuracy even under real-world conditions.
The impressive visual quality and ubiquity of AI-generated images call for reliable and robust detection methods. Reconstruction-based detectors have emerged as a promising direction for transparent and training-free identification of synthetic images. However, due to their fundamentally different mode of operation (compared to standard, classifier-based methods), little is known about their adversarial robustness. In this work, we propose two novel attack methods targeted at detectors that leverage autoencoder reconstruction error. We find that by constructing imperceptible adversarial examples, the distance between original and reconstruction can be artificially increased, causing fake images to be wrongly classified as real. Our evaluation including images from three state-of-the-art generators and three detectors demonstrates that detection performance is significantly decreased, even if attacked images additionally undergo real-world degradations. Critically, our adversarial examples naturally transfer across detectors, as they all share the same principle, pointing towards an inherent vulnerability of reconstruction-based detectors.