ARTFEED — Contemporary Art Intelligence

Activation Watermarking Protects LLMs from Adaptive Attacks

ai-technology · 2026-07-30

Researchers propose Activation Watermarking (AWM), a method to secure LLM monitoring against adaptive attackers who possess local copies and can search offline for prompts that trigger harmful responses while evading detection. AWM randomizes monitoring by fine-tuning the LLM so that its hidden states align with a secret key-derived direction whenever a response violates policy. Detection relies on a similarity test of activations already computed by the provider. Attackers who know everything except the key must optimize against surrogate detectors keyed differently. The approach preserves detection rates for non-adaptive users while resisting adaptive evasion. The paper is available on arXiv under ID 2603.23171.

Key facts

  • arXiv ID: 2603.23171
  • Announce type: replace-cross
  • Proposed method: Activation Watermarking (AWM)
  • AWM uses limited fine-tuning to align hidden states with a secret key-derived direction
  • Detection is a similarity test on activations
  • Adaptive attackers have local copies and can search offline
  • Attackers must optimize against differently keyed surrogate detectors
  • AWM maintains detection rates for non-adaptive users

Entities

Institutions

  • arXiv

Sources