GAP-URGENet Wins ICASSP 2026 URGENT Challenge with Generative-Predictive Fusion
A new framework named GAP-URGENet has been developed by researchers, designed for universal speech enhancement through generative-predictive fusion. It excelled in the blind-test phase of Track 1 at the ICASSP 2026 URGENT Challenge, securing a top position. The framework consists of two interconnected branches: one generative branch focuses on comprehensive speech restoration within a self-supervised representation domain and reconstructs the waveform using a neural vocoder, while the other predictive branch enhances the spectrogram domain. A post-processing module merges the outputs from both branches and extends bandwidth to produce an enhanced waveform at 48 kHz, which is subsequently downsampled. This innovative approach enhances robustness and perceptual quality, addressing the complexities of universal speech enhancement. The research paper can be found on arXiv with the identifier 2604.01832, and audio samples are included. This work marks a notable leap in speech processing, with implications for hearing aids, telecommunications, and voice assistants.
Key facts
- GAP-URGENet is a generative-predictive fusion framework for universal speech enhancement.
- It was developed for Track 1 of the ICASSP 2026 URGENT Challenge.
- The system integrates a generative branch and a predictive branch.
- The generative branch performs full-stack speech restoration in a self-supervised representation domain.
- The predictive branch performs spectrogram-domain enhancement.
- A post-processing module fuses outputs and performs bandwidth extension to 48 kHz.
- The system achieved top performance in the blind-test phase and ranked 1st in objective evaluation.
- The paper is available on arXiv (ID: 2604.01832) with audio examples.
Entities
Institutions
- ICASSP
- URGENT Challenge
- arXiv