DECAF: Decomposing Perturbation Responses into Evidence, Contradiction, and Fragility
A recent paper on arXiv (2608.12935) presents DECAF (Decomposition of Evidence, Contradiction, And Fragility), a novel approach for interpreting model decisions by breaking down perturbation responses into three distinct elements: evidence (E), contradiction (C), and fragility (F). This method overcomes the drawback that the magnitude of responses alone fails to convey the significance of a model's reaction. DECAF analyzes how the differences between factual and counterfactual inputs evolve as paired inputs are gradually disclosed, utilizing the final contrast for trajectory interpretation. It maintains the ordinary magnitude precisely, adhering to Abs = E + C + F, and is distinctive under endpoint-relative axioms. The paper demonstrates DECAF's effectiveness in controlled vision and tabular environments, revealing that the components align with independently measured behaviors. Additionally, a 72-model audit of ImageNet-9 is performed, comparing the components across various models, making this work pertinent to explainable AI and model interpretability by providing a deeper insight into perturbation-based explanations.
Key facts
- Paper arXiv:2608.12935 introduces DECAF (Decomposition of Evidence, Contradiction, And Fragility).
- DECAF decomposes perturbation responses into evidence (E), contradiction (C), and fragility (F).
- The decomposition preserves ordinary magnitude exactly: Abs = E + C + F.
- DECAF is unique under endpoint-relative axioms.
- The method tracks how contrast develops as paired inputs are progressively revealed.
- Validation was performed on controlled vision and tabular settings.
- A 72-model ImageNet-9 audit was conducted.
- The paper addresses limitations of perturbation methods in explainable AI.
Entities
Institutions
- arXiv