ARTFEED — Contemporary Art Intelligence

GhostVAE: Stealthy Backdoor Attack Evades Diffusion Semantic Watermarks

ai-technology · 2026-08-04

Researchers have unveiled GhostVAE, an innovative backdoor attack aimed at undermining semantic watermarking in Latent Diffusion Models (LDMs). This method takes advantage of the neural network-based watermark detection system by embedding a covert backdoor within the encoder of a Variational Autoencoder (VAE), allowing it to bypass watermark detection effectively. GhostVAE functions in two phases: initially, it creates a universal trigger through power spectrum regularization to bolster trigger resilience; subsequently, it trains a compromised VAE encoder using a parameter-aligned goal. Comprehensive assessments across three leading semantic watermarking techniques and three popular LDMs reveal that GhostVAE maintains a watermark detection success rate of 94.4% on non-malicious images while facilitating successful evasion, exposing a significant weakness in existing watermarking methods and highlighting the necessity for stronger protections in AI-generated materials.

Key facts

  • GhostVAE is a backdoor attack targeting semantic watermarking in Latent Diffusion Models (LDMs).
  • The attack plants a backdoor into the VAE encoder to evade watermark detection.
  • GhostVAE uses power spectrum regularization to construct a universal trigger.
  • It trains a backdoored VAE encoder with a parameter-aligned objective.
  • Evaluations were conducted on three semantic watermarking schemes and three LDMs.
  • GhostVAE achieves an average true positive rate of 94.4% on benign images.
  • The attack enables reliable evasion of watermark detection.
  • The research highlights vulnerabilities in neural network-based watermark detection pipelines.

Entities

Institutions

  • arXiv

Sources