Query-Based Transformers Enhance sEMG Hand Gesture Recognition for Prosthetics
A new research paper proposes a query-based transformer architecture for multimodal surface electromyography (sEMG) hand gesture recognition in prosthetic control. Current deep learning approaches, the gold standard, struggle to scale as the number of hand movements increases due to the statistical complexity of decoding expanded gesture sets. State-of-the-art methods rely on low-latency unimodal convolutional architectures, which operate locally and cannot capture long-range sequential patterns. Unimodal setups also fail to leverage complementary information from coordinated signals like inertial and eye-tracking data. The proposed architecture integrates local and global features across multimodal physiological sequences to address these limitations. The paper is available on arXiv under ID 2607.22779.
Key facts
- arXiv ID: 2607.22779
- Announce Type: cross
- Focus on hand gesture recognition via sEMG for prosthetic control
- Deep learning is current gold standard
- Performance degrades as number of hand movements increases
- Current architectures primarily use low-latency unimodal convolutional methods
- Convolutions limit ability to capture long-range sequential patterns
- Proposed solution: query-based transformers integrating local and global features across multimodal sequences
Entities
Institutions
- arXiv