I2VShield: Proactive Defense Against DiT-Based Image-to-Video Models
Researchers propose I2VShield, a generative adversarial attack method to protect privacy against Diffusion Transformer (DiT)-based image-to-video (I2V) models. Current proactive defenses rely on gradient-based attacks requiring high GPU memory, but I2VShield reduces computational overhead via a text-adaptive perturbation framework and an untargeted Multimodal Attention Disruption (MAD) attack. The method aims to prevent misuse of I2V models while maintaining visual imperceptibility.
Key facts
- I2VShield is a proactive defense against DiT-based I2V models
- It uses generative adversarial attacks instead of gradient-based methods
- Includes a text-adaptive perturbation generation framework
- Includes an untargeted Multimodal Attention Disruption (MAD) attack
- Reduces computational overhead and GPU memory requirements
- Maintains visual imperceptibility of perturbations
- Addresses misuse of video generation models
- Published on arXiv with ID 2607.25522
Entities
Institutions
- arXiv