SAPER: Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers
A new paper on arXiv (2608.00264) introduces SAPER (Soft Attention PrunER), an end-to-end differentiable pruning framework for Vision Transformers (ViTs) that leverages interpretability insights to reduce computational costs while preserving accuracy. The authors propose a spectral analysis and visualization technique based on Laplacian eigenvectors of attention maps to understand individual attention heads. They cluster heads semantically to identify functional redundancies, then apply the LapSum Soft Top-K approach for soft pruning. Experiments on ImageNet-1K show SAPER achieves a favorable accuracy-efficiency trade-off, outperforming the RAPTOR baseline in FLOPs reduction. The work addresses the computational demands of large vision foundation models like DINOv2, offering a more efficient and interpretable solution.
Key facts
- Paper arXiv:2608.00264 introduces SAPER (Soft Attention PrunER).
- SAPER is an end-to-end differentiable pruning framework for Vision Transformers.
- The method uses spectral analysis and Laplacian eigenvectors of attention maps.
- Attention heads are semantically clustered to identify functional redundancies.
- LapSum Soft Top-K approach is used for soft pruning.
- Experiments on ImageNet-1K demonstrate favorable accuracy-efficiency trade-off.
- SAPER outperforms RAPTOR baseline in FLOPs reduction.
- The work targets computational demands of vision foundation models like DINOv2.
Entities
Institutions
- arXiv