GhostVAE: Stealthy Backdoor Attack Evades Diffusion Semantic Watermarks
Researchers have unveiled GhostVAE, an innovative backdoor attack aimed at undermining semantic watermarking in Latent Diffusion Models (LDMs). This method takes advantage of the neural network-based watermark detection system by embedding a covert backdoor within the encoder of a Variational Autoencoder (VAE), allowing it to bypass watermark detection effectively. GhostVAE functions in two phases: initially, it creates a universal trigger through power spectrum regularization to bolster trigger resilience; subsequently, it trains a compromised VAE encoder using a parameter-aligned goal. Comprehensive assessments across three leading semantic watermarking techniques and three popular LDMs reveal that GhostVAE maintains a watermark detection success rate of 94.4% on non-malicious images while facilitating successful evasion, exposing a significant weakness in existing watermarking methods and highlighting the necessity for stronger protections in AI-generated materials.
Key facts
- GhostVAE is a backdoor attack targeting semantic watermarking in Latent Diffusion Models (LDMs).
- The attack plants a backdoor into the VAE encoder to evade watermark detection.
- GhostVAE uses power spectrum regularization to construct a universal trigger.
- It trains a backdoored VAE encoder with a parameter-aligned objective.
- Evaluations were conducted on three semantic watermarking schemes and three LDMs.
- GhostVAE achieves an average true positive rate of 94.4% on benign images.
- The attack enables reliable evasion of watermark detection.
- The research highlights vulnerabilities in neural network-based watermark detection pipelines.
Entities
Institutions
- arXiv