PASE: Phonologically Anchored Speech Enhancer Reduces Hallucinations in Generative SE
A new generative speech enhancement model, Phonologically Anchored Speech Enhancer (PASE), has been proposed to address hallucination issues in severe noise conditions. The research, released on arXiv (2511.13300), identifies two types of hallucinations: linguistic (incorrect spoken content) and acoustic (inconsistent speaker characteristics). The authors argue that linguistic hallucination is more fundamental, stemming from models' failure to constrain valid phonological structures. Existing approaches using language models (LMs) are limited because they learn from noise-corrupted representations, leading to contaminated priors. PASE leverages the phonological prior of WavLM to anchor the generation process, reducing hallucinations and improving perceptual quality. The paper is a cross-type announcement, indicating it may have been presented elsewhere. The work is significant for the field of speech enhancement, as it targets a critical limitation of generative models in real-world noisy environments.
Key facts
- PASE is a generative speech enhancement model.
- It addresses linguistic and acoustic hallucinations.
- Linguistic hallucination is considered more fundamental.
- Existing LM-based approaches suffer from contaminated priors.
- PASE leverages the phonological prior of WavLM.
- The paper is available on arXiv with ID 2511.13300.
- The announcement type is cross.
- The model aims to improve perceptual quality.
Entities
Institutions
- arXiv