ARTFEED — Contemporary Art Intelligence

Resource-Aware Mixture of Speech Encoders for Multilingual S2TT

ai-technology · 2026-08-06

A newly proposed framework, MSRT, detailed in a paper on arXiv (2608.04586), tackles the challenges posed by multilinguality in many-to-many speech-to-text translation (S2TT) for multimodal large language models (MLLMs). The issue stems from the use of a single speech encoder across various languages, leading to dominance by high-resource languages and a decline in performance for low-resource ones. MSRT features a resource-aware Mixture of Speech Encoders (MoSE) along with a dedicated language router that allocates each utterance to the suitable expert encoder. It includes a frozen expert for high-resource languages and a trainable one for medium- and low-resource languages. Additionally, a five-stage curriculum learning strategy is introduced to lessen data needs. The authors of the paper are researchers, and it can be accessed on arXiv.

Key facts

  • MSRT is a novel framework for many-to-many speech-to-text translation.
  • It uses a resource-aware Mixture of Speech Encoders (MoSE).
  • MoSE includes an explicit language router.
  • A frozen expert encoder preserves high-resource language capabilities.
  • A trainable expert encoder adapts to medium- and low-resource languages.
  • The framework introduces a five-stage curriculum learning strategy.
  • The paper is available on arXiv with ID 2608.04586.
  • The approach aims to improve multilingual consistency in S2TT.

Entities

Institutions

  • arXiv

Sources