Search papers, labs, and topics across Lattice.
This paper introduces I2VShield, a novel proactive defense framework designed to counteract the misuse of image-to-video (I2V) models, particularly those based on Diffusion Transformers (DiT). By integrating a text-adaptive perturbation generation framework with an untargeted Multimodal Attention Disruption (MAD) attack, I2VShield effectively reduces the computational overhead typically associated with adversarial defenses while maintaining visual imperceptibility. Experimental results show that I2VShield significantly enhances protection against I2V models, particularly in disrupting spatiotemporal coherence, while requiring less GPU memory than existing methods.
I2VShield disrupts DiT-based image-to-video models with minimal computational cost, achieving high protection performance without the need for extensive GPU resources.
The rapid advancement of video generation models has led to the increasing misuse of image-to-video (I2V) models. Although substantial progress has been made in detecting AI-generated videos, proactive defenses against I2V models remain underexplored. In particular, current proactive defenses against I2V models predominantly rely on gradient-based adversarial attacks, which require defenders to possess GPUs with substantial memory resources (VRAM) to generate adversarial examples. To address this issue, we propose I2VShield, a privacy protection method based on generative adversarial attacks tailored to Diffusion Transformer (DiT)-based I2V models. The proposed method primarily consists of two components: (1) a text-adaptive perturbation generation framework integrating adversarial learning to mitigate computational overhead while maintaining visual imperceptibility; and (2) an untargeted Multimodal Attention Disruption (MAD) attack that exploits the inherent vulnerabilities of DiT-based I2V models, maximizing the deviation of the internal attention features from their clean states. Extensive experiments demonstrate that our approach achieves highly competitive protection performance across various datasets and mainstream DiT-based I2V models, particularly in disrupting spatiotemporal coherence, while substantially reducing computational costs.