Deep Learning Enhances Relative Transfer Matrix Estimation for Speech Enhancement
A recent paper on arXiv (2608.11627) presents innovative deep learning techniques aimed at estimating the Relative Transfer Matrix (ReTM), which serves as an extension of the relative transfer function applicable to multiple sources and receivers. This ReTM exhibits potential for improving speech clarity in noisy settings. The researchers propose three supervised learning approaches: convolutional networks in the time domain and short-time frequency transform domain, along with a Long Short-Term Memory (LSTM) recurrent neural network. Experimental results indicate that these models provide superior ReTM estimation compared to the traditional covariance-based method, evaluated through five objective metrics. The study also confirms the frameworks' effectiveness in enhancing speech quality. The full paper can be accessed at https://arxiv.org/abs/2608.11627.
Key facts
- Paper arXiv:2608.11627 proposes deep learning-based ReTM estimation.
- Three frameworks: time-domain CNN, STFT-domain CNN, and LSTM-based RNN.
- Models outperform covariance-based method on five objective metrics.
- ReTM generalizes relative transfer function for multiple receivers and sources.
- ReTM estimation exploits covariance matrices of multichannel recordings.
- Proposed models show effectiveness for speech enhancement.
- Paper type: cross (announcement).
- Published on arXiv.
Entities
Institutions
- arXiv