MEUSLI: Open-Source Multilingual Speech AI for 28 European Languages
A team of researchers has unveiled MEUSLI, a groundbreaking family of open-science multilingual projectors that integrates a Whisper speech encoder with open-source multilingual large language models (LLMs). This innovation facilitates fully open-source end-to-end automatic speech recognition (ASR) across 28 European languages, encompassing both high- and low-resource languages. MEUSLI builds upon previous monolingual systems and employs continual learning methods to adapt seamlessly to languages not included in its training. In addition to ASR, this projector can also be utilized for multilingual speech translation and topic identification with minimal task-specific supervision. The findings are published in a paper on arXiv (2607.22100), marking a pivotal advancement towards inclusive, multilingual AI systems free from proprietary models and English-centric limitations.
Key facts
- MEUSLI is the first open-science multilingual projector family.
- It links a Whisper encoder with open-source multilingual LLMs.
- Supports end-to-end ASR in 28 European languages.
- Uses continual learning to extend to unseen languages.
- Enables speech translation and topic identification with minimal supervision.
- Paper available on arXiv (2607.22100).
- Aims to reduce reliance on proprietary models.
- Covers both high- and low-resource languages.
Entities
Institutions
- arXiv