Adversarial Examples That Fool Humans but Not AI Models
A recent investigation published on arXiv (2607.22722) delves into adversarial examples that are easily identifiable by humans but do not mislead AI systems. Unlike conventional adversarial strategies that introduce subtle alterations to induce errors, this study focuses on significant, conspicuous modifications that maintain the model's accurate predictions while perplexing human observers. The researchers address three key questions: whether humans struggle more than the model with these examples, if standard out-of-distribution (OOD) detection and calibration techniques identify them, and whether current defenses are effective against them. Tests conducted on MNIST, CIFAR-10, and ImageNet reveal that an independent recognizer proxy's accuracy plummets to approximately 49% on CIFAR-10, while the model achieves 100%. This discrepancy is supported by a small human pilot study (N=5) and is not attributed to signal loss, as a matched-magnitude Gaussian control reduces recognizability more rapidly. A CLIP zero-shot proxy further validates this gap at the ImageNet scale, and the research also evaluates confidence and energy-based OOD detection approaches.
Key facts
- arXiv paper 2607.22722 studies adversarial examples with large visible perturbations.
- Unlike typical attacks, these perturbations cause models to keep correct predictions while humans fail to recognize images.
- Three questions tested: human vs model performance, OOD detection, and defense mitigation.
- Experiments conducted on MNIST, CIFAR-10, and ImageNet datasets.
- Recognizer proxy accuracy drops to ~49% on CIFAR-10, model stays at 100%.
- Human pilot (N=5) corroborates the performance gap.
- Gaussian control shows signal loss does not explain the gap.
- CLIP zero-shot proxy confirms the gap on ImageNet.
Entities
Institutions
- arXiv