Supervised and Unsupervised Speech Enhancement Using Nonnegative Matrix Factorization
arXiv:1709.05362 · doi:10.1109/TASL.2013.2270369
Abstract
Reducing the interference noise in a monaural noisy speech signal has been a challenging task for many years. Compared to traditional unsupervised speech enhancement methods, e.g., Wiener filtering, supervised approaches, such as algorithms based on hidden Markov models (HMM), lead to higher-quality enhanced speech signals. However, the main practical difficulty of these approaches is that for each noise type a model is required to be trained a priori. In this paper, we investigate a new class of supervised speech denoising algorithms using nonnegative matrix factorization (NMF). We propose a novel speech enhancement method that is based on a Bayesian formulation of NMF (BNMF). To circumvent the mismatch problem between the training and testing stages, we propose two solutions. First, we use an HMM in combination with BNMF (BNMF-HMM) to derive a minimum mean square error (MMSE) estimator for the speech signal with no information about the underlying noise type. Second, we suggest a scheme to learn the required noise BNMF model online, which is then used to develop an unsupervised speech enhancement system. Extensive experiments are carried out to investigate the performance of the proposed methods under different conditions. Moreover, we compare the performance of the developed algorithms with state-of-the-art speech enhancement schemes using various objective measures. Our simulations show that the proposed BNMF-based methods outperform the competing algorithms substantially.
References in corpus (1)
Cited by in corpus (20)
- Speech enhancement with variational autoencoders and alpha-stable distributions
- ClearBuds: Wireless Binaural Earbuds for Learning-Based Speech Enhancement
- Multi-Objective Learning and Mask-Based Post-Processing for Deep Neural Network Based Speech Enhancement
- Speech Dereverberation Using Nonnegative Convolutive Transfer Function and Spectro temporal Modeling
- SkipConvGAN: Monaural Speech Dereverberation using Generative Adversarial Networks via Complex Time-Frequency Masking
- A State-Space Approach to Dynamic Nonnegative Matrix Factorization
- Unsupervised Low Latency Speech Enhancement with RT-GCC-NMF
- Low Rank and Sparsity Analysis Applied to Speech Enhancement via Online Estimated Dictionary
- Deep Speech Enhancement for Reverberated and Noisy Signals using Wide Residual Networks
- Modeling the Comb Filter Effect and Interaural Coherence for Binaural Source Separation
- Speech enhancement with weakly labelled data from AudioSet
- Normalized Features for Improving the Generalization of DNN Based Speech Enhancement
- Adaptive dictionary based approach for background noise and speaker classification and subsequent source separation
- End-to-end speech enhancement based on discrete cosine transform
- A Speech Enhancement Algorithm based on Non-negative Hidden Markov Model and Kullback-Leibler Divergence
- Improving the Intelligibility of Electric and Acoustic Stimulation Speech Using Fully Convolutional Networks Based Speech Enhancement
- Semi-supervised Speech Enhancement in Envelop and Details Subspaces
- Directional Embedding Based Semi-supervised Framework For Bird Vocalization Segmentation
- PROSE: Perceptual Risk Optimization for Speech Enhancement
- Single Channel Speech Enhancement Using Temporal Convolutional Recurrent Neural Networks