ARTFEED — Contemporary Art Intelligence

CAM Review: 57 Papers on Visual Explanations in CNN, Transformer, and Foundation Models

publication · 2026-08-13

A synthesis of 57 studies published since 2016, accessible on arXiv (2608.12299), evaluates class activation mapping (CAM). This review outlines the progression of CAM, starting from global-average-pooled CNN classifiers and extending to various methodologies such as gradient-based explanations, high-resolution upscaling, weakly supervised localization, transformer token attribution, and foundation-model strategies utilizing CLIP, DINO, and SAM. It introduces a taxonomy centered on attribution mechanisms to elucidate visual explanation methods. The focus is on CAM's role in transforming internal model data into heatmaps that emphasize significant areas of images. Notably, the review is centered on methodologies and does not include experimental comparisons or benchmarks, catering to an academic audience in the fields of computer vision and explainable AI.

Key facts

  • The review synthesizes 57 method-centered papers on class activation mapping (CAM).
  • Papers were published from 2016 onward.
  • The review is available on arXiv with ID 2608.12299.
  • It is announced as a cross preprint.
  • CAM converts internal model evidence into heatmaps highlighting image regions, channels, tokens, or patches.
  • Methods covered include gradient-based, gradient-free, high-resolution upscaling, weakly supervised localization, transformer token attribution, causal and debiasing, and foundation-model approaches.
  • Foundation-model approaches use CLIP, DINO, SAM, or feature-distribution comparisons.
  • The paper develops a taxonomy separating methods by attribution mechanism.

Entities

Institutions

  • arXiv

Sources