Whisper-Large-v3-Turbo and IndicWav2Vec Tie in Nepali ASR Benchmark
A new research paper on arXiv (2608.12327) presents the first thorough analysis of six multilingual pre-trained models aimed at Nepali automatic speech recognition (ASR). The models examined include XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium, Whisper-Large-v3-Turbo, and Conformer-Hi, showcasing different architectures like CTC self-supervised and hybrid Conformer-CTC. All the models were fine-tuned using the OpenSLR SLR54 Nepali dataset, which has about 165 hours of audio, ensuring consistent preprocessing and learning parameters. They were tested across three separate datasets focusing on metrics like Word Error Rate (WER) and Character Error Rate (CER). Interestingly, Whisper-Large-v3-Turbo and IndicWav2Vec performed similarly, suggesting that related language families can balance out differences in model size for Nepali ASR. The study also notes that CTC decoders are quicker, making them better for real-time applications. This research fills a void in Nepali ASR studies, setting a benchmark for future work.
Key facts
- First controlled benchmark of six multilingual pre-trained models for Nepali ASR
- Models fine-tuned on OpenSLR SLR54 Nepali corpus (~165 hours)
- Identical preprocessing, splits, optimizer, and learning-rate schedules used
- Evaluated on OpenSLR, FLEURS, and Common Voice test sets
- Whisper-Large-v3-Turbo achieved 14.76% WER
- IndicWav2Vec achieved 14.89% WER
- 9x parameter gap and 40x pretraining-data gap between top models
- CTC decoders run faster, suitable for real-time applications
Entities
Institutions
- arXiv
- OpenSLR
- FLEURS
- Common Voice