Whisper-Aware LLM Cuts Hallucinations and Errors in Whispered Speech Recognition
A new approach named Whisper-Aware LLM has been unveiled to tackle the difficulties associated with automatic speech recognition (ASR) for whispered communication. This system, outlined in a paper on arXiv (2608.10836), enables an Audio-LLM to recognize and respond to the uncertainties present in whispered audio signals. By engaging in self-supervised tasks to assess physical limitations, the model cultivates a form of self-awareness. A unique Confidence-Fused Decoding mechanism is employed to provide the LLM decoder with high-level guidance and frame-level attention adjustments. Testing reveals a 17% relative decrease in Character Error Rate (CER) on the AISHELL6-Whisper benchmark, marking a new state of the art. Importantly, this method also improves reliability, reducing hallucination rates from over 25% to 4.5%, addressing two major ASR failure modes: the inability to capture whispered speech and erroneous transcription of background noise.
Key facts
- The Whisper-Aware LLM framework is introduced for robust whispered speech recognition.
- It teaches an Audio-LLM to perceive and react to uncertainty in whispered speech.
- The model learns to quantify physical deficiencies of acoustic signals via self-supervised tasks.
- A novel Confidence-Fused Decoding mechanism provides high-level instructions and frame-level attention modulation.
- The model achieves a 17% relative CER reduction on AISHELL6-Whisper, setting a new state-of-the-art.
- Hallucination rates drop from over 25% to 4.5%.
- The paper is available on arXiv with identifier 2608.10836.
- The approach addresses two failure modes: failing to capture whispered speech and hallucinatory transcription of noise.
Entities
Institutions
- arXiv