Alibaba AI voice model ranks top 5 globally, beats US rivals on regional accents
Alibaba Group Holding's new AI voice model has entered the top five globally on the Speech Arena benchmark, outperforming Western rivals OpenAI and xAI. The model supports over 30 languages, seven major Chinese dialects, and more than 20 regional accents. The Speech Arena, operated by San Francisco-based Artificial Analysis and backed by Nat Friedman and Andrew Ng, ranks models via blind user evaluations of speech clips using an Elo system. Users assess three core capabilities: speech-to-text, end-to-end voice understanding and conversation, and text-to-speech. In a separate Word Error Rate index by Artificial Analysis, Alibaba's Fun-Realtime-ASR model ranked first with a 1.8% error rate, meaning fewer than two words per 100 were transcribed incorrectly. Alibaba owns the South China Morning Post.
Key facts
- Alibaba AI voice model cracked top 5 globally on Speech Arena benchmark.
- Outperformed Western rivals OpenAI and xAI.
- Supports over 30 languages, seven Chinese dialects, and over 20 regional accents.
- Speech Arena is operated by Artificial Analysis, backed by Nat Friedman and Andrew Ng.
- Benchmark uses blind user evaluations of speech clips with Elo-based ranking.
- Evaluates speech-to-text, voice understanding, and text-to-speech capabilities.
- Alibaba's Fun-Realtime-ASR model ranked first in Word Error Rate index at 1.8%.
- Alibaba owns the South China Morning Post.
Entities
Institutions
- Alibaba Group Holding
- OpenAI
- xAI
- Artificial Analysis
- South China Morning Post
Locations
- San Francisco
- United States
- Beijing
- China