ARTFEED — Contemporary Art Intelligence

StuPASE: Low-Hallucination Studio-Quality Speech Enhancement

ai-technology · 2026-08-06

Researchers have introduced StuPASE, a generative speech enhancement model that achieves studio-level quality while minimizing hallucinations. Built upon the PASE framework, StuPASE addresses two key limitations: it improves dereverberation by finetuning with dry targets instead of those with simulated early reflections, and it replaces the GAN-based module with a flow-matching module to handle strong additive noise. Experiments show StuPASE consistently produces high-quality speech with low hallucination, outperforming state-of-the-art methods. Audio demos are available online.

Key facts

  • StuPASE is a generative speech enhancement model.
  • It is built upon the PASE framework.
  • Finetuning with dry targets improves dereverberation.
  • Flow-matching module replaces GAN-based module for noise robustness.
  • StuPASE achieves studio-level quality with low hallucination.
  • Outperforms state-of-the-art SE methods.
  • Audio demos available at https://xiaobin (truncated).
  • Paper announced on arXiv with ID 2603.09234.

Entities

Institutions

  • arXiv

Sources