ARTFEED — Contemporary Art Intelligence

Ge²mS-T: Ultra-Efficient Spiking Transformer via Multi-Dimensional Grouping

ai-technology · 2026-08-07

A novel architecture for Spiking Vision Transformers (S-ViTs) seeks to enhance energy efficiency significantly while addressing training and inference challenges. The newly proposed model, Ge²mS-T, incorporates grouped computation across temporal, spatial, and network structure dimensions. It utilizes the Grouped-Exponential-Coding-based IF (ExpG-IF) model, which allows for lossless conversion with consistent training overhead and accurate spike pattern management, alongside Group-wise Spiking Self-Attention (GW-SSA) to lower computational complexity through multi-scale token grouping. This research tackles the limitations of current methods such as ANN-SNN Conversion and Spatial-Temporal Backpropagation (STBP), which hinder simultaneous optimization of memory, accuracy, and energy use. The paper can be found on arXiv with the identifier 2604.08894, detailing the methodology. This work holds significance for neuromorphic computing and energy-efficient AI, with potential implications for edge computing and real-time applications.

Key facts

  • Ge²mS-T is a novel architecture for Spiking Vision Transformers (S-ViTs).
  • It implements grouped computation across temporal, spatial, and network structure dimensions.
  • Introduces Grouped-Exponential-Coding-based IF (ExpG-IF) model for lossless conversion.
  • Develops Group-wise Spiking Self-Attention (GW-SSA) to reduce computational complexity.
  • Addresses limitations of ANN-SNN Conversion and Spatial-Temporal Backpropagation (STBP).
  • Aims for concurrent optimization of memory, accuracy, and energy consumption.
  • Paper available on arXiv with identifier 2604.08894.
  • Targets ultra-high energy efficiency in spiking transformers.

Entities

Institutions

  • arXiv

Sources