ARTFEED — Contemporary Art Intelligence

LLM Refusal Asymmetry: Answer Release Local, Suppression Global

ai-technology · 2026-08-18

A recent preprint on arXiv (2608.15772) delves into the inner workings of large language models (LLMs) when they decline to respond to prompts. By employing a controlled withholding framework to align answering and refusal trajectories, the research uncovers a 'broken symmetry' in the locality of interventions. Notably, even when a model issues a clear refusal, its hidden states still allow for the correct answer to be linearly extracted. The act of revealing this withheld answer is a localized task, needing just a single-position adjustment. In contrast, reinstating suppression demands wider interventions across several positions, making the formation of a coherent refusal sequence significantly more challenging. These results imply that refusal involves suppression rather than mere information deletion, raising important considerations for AI safety and interpretability. The study, conducted by a team of researchers, highlights ongoing efforts in AI alignment and model behavior.

Key facts

  • The paper is titled 'Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration'.
  • It is available on arXiv with identifier 2608.15772.
  • The study uses a controlled withhold setting to create matched answering and refusal trajectories.
  • Bidirectional activation patching is used to investigate causal interventions.
  • The correct answer remains linearly recoverable from hidden states even during refusal.
  • Releasing the withheld answer requires only a single-position patch.
  • Reimposing suppression requires broader interventions across multiple positions.
  • The phenomenon is termed 'broken symmetry'.

Entities

Institutions

  • arXiv

Sources