StuPASE: Low-Hallucination Studio-Quality Speech Enhancement
Researchers have introduced StuPASE, a generative speech enhancement model that achieves studio-level quality while minimizing hallucinations. Built upon the PASE framework, StuPASE addresses two key limitations: it improves dereverberation by finetuning with dry targets instead of those with simulated early reflections, and it replaces the GAN-based module with a flow-matching module to handle strong additive noise. Experiments show StuPASE consistently produces high-quality speech with low hallucination, outperforming state-of-the-art methods. Audio demos are available online.
Key facts
- StuPASE is a generative speech enhancement model.
- It is built upon the PASE framework.
- Finetuning with dry targets improves dereverberation.
- Flow-matching module replaces GAN-based module for noise robustness.
- StuPASE achieves studio-level quality with low hallucination.
- Outperforms state-of-the-art SE methods.
- Audio demos available at https://xiaobin (truncated).
- Paper announced on arXiv with ID 2603.09234.
Entities
Institutions
- arXiv