SAGE: Adaptive Surrogate Gradients for Spiking Transformers
A new approach called SAGE has been developed by researchers for training spiking neural networks (SNNs), enhancing accuracy by adapting surrogate gradients throughout the training process. This innovative technique, outlined in a paper available on arXiv (2608.13702), tackles the optimization difficulties associated with SNNs, which are known for their energy efficiency but struggle due to the non-differentiable nature of the spike function. SAGE leverages block-level uncertainty derived from normalized self-attention entropy to modify the surrogate-gradient slope during training, while keeping the inference model intact. Tests conducted on CIFAR-10/100 demonstrate that SAGE outperforms fixed-surrogate benchmarks, maintaining the original architecture and deployment costs while offering greater optimization flexibility.
Key facts
- SAGE is an uncertainty-modulated surrogate-gradient mechanism for Transformer-based SNNs.
- It estimates block-level uncertainty from normalized self-attention entropy.
- The method adapts the surrogate-gradient slope during training only.
- Inference model remains unchanged, preserving deployment cost.
- Experiments on CIFAR-10/100 show improved accuracy over fixed-surrogate baselines.
- The paper is available on arXiv with ID 2608.13702.
- SNNs offer energy-efficient alternative to conventional deep neural networks.
- Training SNNs is challenging due to non-differentiable spike function.
Entities
Institutions
- arXiv