TransSLR: Lightweight Transformer for Sign Language Recognition
A new lightweight transformer model, TransSLR, has been developed to address the challenge of automated sign language recognition for under-represented languages, specifically Central African Sign Language (CASL). The model, trained from scratch on 64-frame normalized pose sequences, achieves signer-independent generalization without relying on visual appearance. The research highlights the failure of fine-tuning high-resource models due to the limited scale of CASL data and the significant domain gap between CASL and large-scale corpora like WLASL. The best reported accuracy on the CASL-W60 benchmark is 69.93%, and TransSLR aims to improve upon this. The study is available on arXiv with the identifier 2608.06407.
Key facts
- TransSLR is a lightweight Temporal Transformer Encoder.
- It is trained on 64-frame normalized pose sequences.
- It uses average pooling and a classification head.
- It operates on geometric keypoint representations rather than raw RGB.
- It targets Central African Sign Language (CASL).
- The CASL-W60 benchmark has a best reported accuracy of 69.93%.
- Fine-tuning high-resource models fails to close the accuracy gap.
- The model is trained from scratch to avoid reliance on pre-trained representations.
Entities
—