White-Box Evasion Attack Framework Defeats Explainable AI Auditors
A new research paper on arXiv (ID 2608.00566) introduces a white-box, gradient-regularized evasion attack framework that can fool explainable AI auditors. The framework, called 'Crushing the Evidence,' uses a continuous-embedding dual-penalty approach to directly penalize the explanation outputs, thereby concealing algorithmic biases or backdoors. This attack is more potent than previous black-box attacks, which relied on scaffolding out-of-distribution (OOD) detectors. The paper highlights a critical vulnerability in post-hoc model explainers like LIME, SHAP, and Integrated Gradients, which are widely used to audit models in high-stakes domains such as finance, healthcare, and social welfare. While defenses have been developed to neutralize black-box attacks by identifying anomalous perturbation footprints, this new white-box attack bypasses those defenses. The research underscores the need for more robust explainability methods to ensure transparency and accountability in AI systems.
Key facts
- Paper ID: arXiv:2608.00566
- Announce Type: cross
- Introduces a white-box, gradient-regularized evasion attack framework
- Uses a continuous-embedding dual-penalty framework
- Targets post-hoc explainers: LIME, SHAP, Integrated Gradients
- Attack can conceal algorithmic biases or backdoors
- Previous black-box attacks relied on OOD detectors
- Defenses exist for black-box attacks but not for this white-box attack
Entities
Institutions
- arXiv