LLM Refusal Asymmetry: Answer Release Local, Suppression Global
A recent preprint on arXiv (2608.15772) delves into the inner workings of large language models (LLMs) when they decline to respond to prompts. By employing a controlled withholding framework to align answering and refusal trajectories, the research uncovers a 'broken symmetry' in the locality of interventions. Notably, even when a model issues a clear refusal, its hidden states still allow for the correct answer to be linearly extracted. The act of revealing this withheld answer is a localized task, needing just a single-position adjustment. In contrast, reinstating suppression demands wider interventions across several positions, making the formation of a coherent refusal sequence significantly more challenging. These results imply that refusal involves suppression rather than mere information deletion, raising important considerations for AI safety and interpretability. The study, conducted by a team of researchers, highlights ongoing efforts in AI alignment and model behavior.
Key facts
- The paper is titled 'Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration'.
- It is available on arXiv with identifier 2608.15772.
- The study uses a controlled withhold setting to create matched answering and refusal trajectories.
- Bidirectional activation patching is used to investigate causal interventions.
- The correct answer remains linearly recoverable from hidden states even during refusal.
- Releasing the withheld answer requires only a single-position patch.
- Reimposing suppression requires broader interventions across multiple positions.
- The phenomenon is termed 'broken symmetry'.
Entities
Institutions
- arXiv