EEGAlign: Joint Text-Audio Alignment for Chinese EEG-to-Text Decoding
Researchers introduced EEGAlign, a parameter-efficient framework that jointly aligns electroencephalography (EEG) signals with both text semantics and audio acoustic features for decoding Chinese speech. The approach addresses challenges in non-invasive neural communication for individuals with severe speech and motor impairments. Unlike existing methods that rely on a single supervisory axis, EEGAlign combines sentence-level discriminability and fine-grained temporal resolution to handle the high-dimensional output space of thousands of Chinese characters, inter-subject variability, and low signal-to-noise ratios. The framework is designed for both speech production and perception decoding from scalp EEG, offering a safer and more deployable alternative to invasive electrocorticography. The study is detailed in a preprint on arXiv (2607.25626).
Key facts
- EEGAlign jointly aligns EEG with text semantics and audio acoustic features.
- It decodes Chinese speech from scalp EEG for non-invasive neural communication.
- Addresses challenges of high-dimensional output space, inter-subject variability, and low SNR.
- Existing methods use single supervisory axis, insufficient for large-vocabulary Chinese decoding.
- Framework is parameter-efficient.
- Targets individuals with severe speech and motor impairments.
- Scalp EEG is safer and more widely deployable than invasive electrocorticography.
- Preprint available on arXiv (2607.25626).
Entities
Institutions
- arXiv