Demystifying Adversarial Robustness in Diffusion Models: Compression, Randomness, and Geometry
A recent paper on arXiv (2505.22839) delves into how diffusion models enhance adversarial robustness in deep neural networks. The researchers found that these models actually increase the ℓp distance from clean samples, challenging the notion that purification brings perturbed images closer to their clean counterparts. They introduce a cohesive explanation that attributes the robustness enhancement to two factors: gradient masking caused by internal randomness and the compression of the image space. The randomness of the model significantly affects the purified images, resulting in gradient masking that existing techniques cannot eliminate. Additionally, the study examines the geometric properties of diffusion-based purification, offering a robust theoretical framework. This work aids in understanding and potentially bolstering adversarial robustness in AI, impacting security and reliability in machine learning.
Key facts
- Paper arXiv:2505.22839v2, announce type replace-cross
- Diffusion models improve empirical adversarial robustness of deep neural networks
- Diffusion models increase ℓp distance to clean samples, rejecting the purification denoising hypothesis
- Robustness improvement decomposed into gradient masking from randomness and compression of image space
- Purified images heavily influenced by internal randomness of diffusion models
- Gradient masking cannot be removed by previous methods
- Paper provides a unifying account of diffusion-based purification
- Research aims to demystify mechanisms underlying diffusion-based robustness
Entities
Institutions
- arXiv