HybridSB-MoE: Dual-Domain Speech Enhancement with Scene-Adaptive Expert Routing
A new generative speech enhancement framework called HybridSB-MoE has been released in a preprint on arXiv (ID 2608.12715). This model tackles shortcomings found in current approaches: spectral models interfere with phase, waveform models overlook harmonics, and Schrödinger Bridges (SB) have loosely connected inference costs. HybridSB-MoE employs asymmetric uncertainty fusion, where the spectral pathway captures epistemic uncertainty and the waveform bridge accounts for aleatoric variance. It also integrates a heterogeneous mixture-of-experts (MoE) with top-k=2 routing across five different architectural types. The goal of this framework is to enhance speech quality by utilizing both spectral and waveform domains, adapting to various acoustic environments. The research is pertinent to audio processing, machine learning, and artificial intelligence.
Key facts
- HybridSB-MoE is a new framework for generative speech enhancement.
- It addresses three gaps: spectral models disrupt phase, waveform models miss harmonics, and Schrödinger Bridges have inference cost not tied to training.
- The framework uses dual-domain processing (spectral and waveform).
- It employs asymmetric uncertainty fusion combining epistemic and aleatoric uncertainty.
- It uses heterogeneous mixture-of-experts with top-k=2 routing across five architectural archetypes.
- The paper is available on arXiv with ID 2608.12715.
- The submission type is cross.
- The framework aims to improve speech enhancement by adapting to distinct error regimes.
Entities
Institutions
- arXiv