MLP Layers and Mid-Network Blocks Encode Refusal Behavior in LLMs
A new study on arXiv (2608.11583) investigates where safety-aligned refusal behavior is encoded in large language models (LLMs). The research, which uses two open-weight model pairs and four safety benchmarks, reveals that refusal transfer is dominated by MLP (multi-layer perceptron) weights rather than attention weights. Replacing MLP parameters recovers substantially more malicious-prompt refusal, with gains of at least 2.7 times across benchmarks. The study also finds a consistent mid-network concentration of refusal-relevant parameters within the MLP stack. This work challenges the assumption that safety alignment is distributed across the entire network, suggesting that it is concentrated in specific layers. The findings have implications for improving LLM safety and understanding model internals.
Key facts
- Study on arXiv: 2608.11583
- Uses two open-weight model pairs and four safety benchmarks
- Refusal transfer dominated by MLP weights
- Replacing MLP parameters recovers at least 2.7 times more refusal than attention
- Refusal-relevant parameters concentrated in mid-network blocks
- Challenges distributed safety alignment assumption
- Implications for LLM safety and interpretability
Entities
Institutions
- arXiv