Fool's Gold: Decoy Hardening Defense Against Safety-Removal Attacks on Open-Weight AI Models
A novel defensive strategy known as 'Fool's Gold' (decoy hardening) tackles vulnerabilities in safety alignment within open-weight language models. According to preprint arXiv:2608.17202v1, current safety alignment can be easily compromised through a process called abliteration. Fool's Gold undermines the refusal strip and contaminates its rewards, generating misleading responses to dangerous prompts that seem plausible but are actually fabricated. This method was evaluated on seven models across five families, ranging from 9B to 122B in size. Six of these models successfully met efficacy criteria, with decoys making up 51% to 90% of their responses, enhancing effectiveness by 27 to 84 percentage points. This research marks a shift in AI safety towards deception instead of merely preserving alignment. The preprint has not undergone peer review yet.
Key facts
- Safety alignment in open-weight language models is trivially removable via abliteration.
- No release-time defense durably prevents abliteration.
- Fool's Gold is a decoy hardening defense that concedes the refusal strip.
- Once refusal is stripped, most answers to hazardous requests are decoys with falsified critical elements.
- Decoys are trained inside a differentiable simulation of the attack and express only in the attacked state.
- A refusal pin and benign leash maintain clean-state behavior.
- The defense was tested on seven models from five families, 9B-122B, dense and mixture-of-experts.
- On six models passing the pre-registered efficacy gate, 0.51-0.90 of attacked-state responses were decoys, with +0.27-0.84 attributable to the defense.
Entities
—