IndexTTS 2.5: Enhanced Multilingual Text-to-Speech with Faster Inference
IndexTTS 2.5 has been launched as an advanced version of the zero-shot neural text-to-speech foundation model, enhancing capabilities over its predecessor, IndexTTS 2. This update increases multilingual coverage, improves inference speed, and elevates synthesis quality. The semantic codec frame rate has been cut from 50 Hz to 25 Hz, halving the sequence length. Additionally, the S2M module benefits from an upgraded backbone, transitioning from U-DiT to Zipformer, which optimizes mel-spectrogram generation. A technical report detailing these advancements can be found on arXiv under identifier 2601.03888.
Key facts
- IndexTTS 2.5 is a new version of the zero-shot neural text-to-speech foundation model.
- It builds on IndexTTS 2, which introduced a transformer-based T2S module and a non-autoregressive S2M module.
- IndexTTS 2.5 enhances multilingual coverage, inference speed, and overall synthesis quality.
- The semantic codec frame rate is reduced from 50 Hz to 25 Hz, halving sequence length.
- The S2M module's backbone is upgraded from U-DiT to Zipformer, reducing parameters and speeding up mel-spectrogram generation.
- The technical report is available on arXiv with identifier 2601.03888.
- The announcement type is 'replace-cross'.
- The model aims to lower training and inference costs while improving performance.
Entities
Institutions
- arXiv