ARTFEED — Contemporary Art Intelligence

White-Box Evasion Attack Framework Defeats Explainable AI Auditors

ai-technology · 2026-08-04

A new research paper on arXiv (ID 2608.00566) introduces a white-box, gradient-regularized evasion attack framework that can fool explainable AI auditors. The framework, called 'Crushing the Evidence,' uses a continuous-embedding dual-penalty approach to directly penalize the explanation outputs, thereby concealing algorithmic biases or backdoors. This attack is more potent than previous black-box attacks, which relied on scaffolding out-of-distribution (OOD) detectors. The paper highlights a critical vulnerability in post-hoc model explainers like LIME, SHAP, and Integrated Gradients, which are widely used to audit models in high-stakes domains such as finance, healthcare, and social welfare. While defenses have been developed to neutralize black-box attacks by identifying anomalous perturbation footprints, this new white-box attack bypasses those defenses. The research underscores the need for more robust explainability methods to ensure transparency and accountability in AI systems.

Key facts

  • Paper ID: arXiv:2608.00566
  • Announce Type: cross
  • Introduces a white-box, gradient-regularized evasion attack framework
  • Uses a continuous-embedding dual-penalty framework
  • Targets post-hoc explainers: LIME, SHAP, Integrated Gradients
  • Attack can conceal algorithmic biases or backdoors
  • Previous black-box attacks relied on OOD detectors
  • Defenses exist for black-box attacks but not for this white-box attack

Entities

Institutions

  • arXiv

Sources