Activation Watermarking Protects LLMs from Adaptive Attacks
Researchers propose Activation Watermarking (AWM), a method to secure LLM monitoring against adaptive attackers who possess local copies and can search offline for prompts that trigger harmful responses while evading detection. AWM randomizes monitoring by fine-tuning the LLM so that its hidden states align with a secret key-derived direction whenever a response violates policy. Detection relies on a similarity test of activations already computed by the provider. Attackers who know everything except the key must optimize against surrogate detectors keyed differently. The approach preserves detection rates for non-adaptive users while resisting adaptive evasion. The paper is available on arXiv under ID 2603.23171.
Key facts
- arXiv ID: 2603.23171
- Announce type: replace-cross
- Proposed method: Activation Watermarking (AWM)
- AWM uses limited fine-tuning to align hidden states with a secret key-derived direction
- Detection is a similarity test on activations
- Adaptive attackers have local copies and can search offline
- Attackers must optimize against differently keyed surrogate detectors
- AWM maintains detection rates for non-adaptive users
Entities
Institutions
- arXiv