UniPASE: Universal Speech Enhancement with High Fidelity and Low Hallucinations
Researchers have unveiled UniPASE, a generative model aimed at universal speech enhancement (USE) that effectively restores speech signals affected by various distortions at different sampling rates. As outlined in the arXiv paper 2604.14606, this model builds upon the low-hallucination PASE framework. At its core lies DeWavLM-Omni, a representation-level enhancement module that has been fine-tuned from WavLM through knowledge distillation using a comprehensive multi-distortion dataset. It transforms impaired waveforms into clear phonetic representations, achieving strong enhancement with minimal hallucination. An Adapter then produces enhanced acoustic representations for a neural Vocoder, which reconstructs high-fidelity 16-kHz waveforms, subsequently converted to 48 kHz by a PostNet before resampling. This model tackles universal speech enhancement issues and is applicable in hearing aids, telecommunications, and voice assistants.
Key facts
- UniPASE is a generative model for universal speech enhancement (USE).
- It extends the low-hallucination PASE framework.
- Core module is DeWavLM-Omni, fine-tuned from WavLM via knowledge distillation.
- Trained on a large-scale supervised multi-distortion dataset.
- Converts degraded waveforms into clean phonetic representations.
- Adapter generates enhanced acoustic representations.
- Neural Vocoder reconstructs high-fidelity 16-kHz waveforms.
- PostNet converts waveforms to 48 kHz before resampling to original rates.
Entities
—