Concept2Scenario: A Framework to Identify Vulnerable Scenarios in LLMs
A recent paper on arXiv (2607.23496) explores the reasons behind the inconsistent refusal of harmful requests by safety-aligned large language models (LLMs). The researchers reveal that prompts framed within specific scenarios trigger internal directives that effectively lower refusal rates. To build on this finding, they introduce Concept2Scenario, a framework designed for identifying vulnerable scenarios through concept-based attribution. This framework utilizes a sparse autoencoder to create a wide-ranging concept space, links refusal suppression to particular concepts, and converts them into clear scenarios. This research connects empirical red-teaming with mechanistic interpretability, providing a structured approach to uncover and comprehend safety weaknesses in LLMs, independent of merely observed attack results.
Key facts
- arXiv paper 2607.23496
- Title: Do LLMs Know Their Vulnerable Scenarios?
- Safety-aligned LLMs are trained to refuse harmful requests
- Embedding harmful requests in particular scenarios can bypass safeguards
- Existing red-teaming identifies effective scenarios through observed attack outcomes
- Mechanistic interpretability studies have characterized refusal directions and jailbreak-associated features
- Scenario-wrapped prompts activate internal scenario directions
- Causal steering consistently reduces refusal scores
- Concept2Scenario is a concept-based attribution framework
- It uses a sparse autoencoder to instantiate a broad concept space
- It attributes refusal suppression to individual concepts
- It translates concepts into interpretable scenarios
Entities
Institutions
- arXiv