ARTFEED — Contemporary Art Intelligence

New Framework Measures Stability of AI Attribution Methods

ai-technology · 2026-08-06

A recent research article introduces a framework based on distribution to assess the stability of attribution methods (AMs) that aim to clarify black-box models. This framework measures the extent of separability in ranked attribution vectors and determines the highest index where feature ranking is consistent. Additionally, it allows for the comparison of AMs by evaluating the stability of their rankings throughout a dataset. This method offers an additional metric for assessing explainer stability, tackling the challenge of fluctuating attribution scores caused by stochastic elements. The study can be found on arXiv titled 'Measuring Explainer Stability via Attribution Separability'.

Key facts

  • The paper proposes a distribution-based framework to capture the stability of attribution scores.
  • Attribution methods assign importance scores to features and are used to explain black-box models.
  • Most attribution methods can produce variable scores due to stochastic components.
  • The framework allows understanding the degree of separability in the ranked attribution vector.
  • It obtains the largest index for which a feature ranking remains reliable.
  • The framework can compare attribution methods based on the robustness of their rankings across a dataset.
  • Experiments demonstrate how to apply the method to evaluate explainer stability.
  • The paper is titled 'Measuring Explainer Stability via Attribution Separability' and is available on arXiv.

Entities

Institutions

  • arXiv

Sources