Circuit Interpretability Evidence Fails Under Analytic Variation, Study Finds
A new study posted on arXiv (2608.13754) challenges the reliability of mechanistic interpretability as evidence for EU AI Act compliance. The researchers pre-registered a grid of seven analytic axes, each with levels drawn from published implementations, and applied them to GPT-2 small on the indirect object identification task. They mapped discovered circuits through a deterministic claim map to structured Annex IV statements. Across 15,840 pre-registered specifications, 7,561 produced a claim. The derived statement flipped across 73.2% of specification pairs (95% CI 0.725 to 0.738), and the modal claim commanded only 41.1% of the space. The authors conclude that circuit-level interpretability evidence does not survive defensible analytic variation, undermining its use in high-risk AI system documentation. The study was announced as new on arXiv and is available at https://arxiv.org/abs/2608.13754.
Key facts
- Study posted on arXiv with ID 2608.13754
- Focuses on mechanistic interpretability and circuit discovery
- Evaluates evidence for EU AI Act high-risk system documentation
- Uses GPT-2 small and indirect object identification task
- Pre-registered 15,840 specifications across seven analytic axes
- 7,561 specifications produced a claim
- Derived statement flips across 73.2% of specification pairs (95% CI 0.725 to 0.738)
- Modal claim commands 41.1% of the space
Entities
Institutions
- arXiv
- EU