Resource-Aware Mixture of Speech Encoders for Multilingual S2TT
A newly proposed framework, MSRT, detailed in a paper on arXiv (2608.04586), tackles the challenges posed by multilinguality in many-to-many speech-to-text translation (S2TT) for multimodal large language models (MLLMs). The issue stems from the use of a single speech encoder across various languages, leading to dominance by high-resource languages and a decline in performance for low-resource ones. MSRT features a resource-aware Mixture of Speech Encoders (MoSE) along with a dedicated language router that allocates each utterance to the suitable expert encoder. It includes a frozen expert for high-resource languages and a trainable one for medium- and low-resource languages. Additionally, a five-stage curriculum learning strategy is introduced to lessen data needs. The authors of the paper are researchers, and it can be accessed on arXiv.
Key facts
- MSRT is a novel framework for many-to-many speech-to-text translation.
- It uses a resource-aware Mixture of Speech Encoders (MoSE).
- MoSE includes an explicit language router.
- A frozen expert encoder preserves high-resource language capabilities.
- A trainable expert encoder adapts to medium- and low-resource languages.
- The framework introduces a five-stage curriculum learning strategy.
- The paper is available on arXiv with ID 2608.04586.
- The approach aims to improve multilingual consistency in S2TT.
Entities
Institutions
- arXiv