Statistical Speech Enhancement Based on Probabilistic Integration of Variational Autoencoder and Non-Negative Matrix Factorization
arXiv:1710.11439 · doi:10.1109/ICASSP.2018.8461530
Abstract
This paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach to speech enhancement is to train a deep neural network (DNN) to take noisy speech as input and output clean speech. Although this supervised approach requires a very large amount of pair data for training, it is not robust against unknown environments. Another approach is to use non-negative matrix factorization (NMF) based on basis spectra trained on clean speech in advance and those adapted to noise on the fly. This semi-supervised approach, however, causes considerable signal distortion in enhanced speech due to the unrealistic assumption that speech spectrograms are linear combinations of the basis spectra. Replacing the poor linear generative model of clean speech in NMF with a VAE---a powerful nonlinear deep generative model---trained on clean speech, we formulate a unified probabilistic generative model of noisy speech. Given noisy speech as observed data, we can sample clean speech from its posterior distribution. The proposed method outperformed the conventional DNN-based method in unseen noisy environments.
5 pages, 3 figures, version that Eqs. (9), (19), and (20) in v2 (submitted to ICASSP 2018) are corrected. Samples available here: http://sap.ist.i.kyoto-u.ac.jp/members/yoshiaki/demo/vae-nmf/
References in corpus (4)
Cited by in corpus (31)
- Speech Enhancement and Dereverberation with Diffusion-based Generative Models
- Dynamical Variational Autoencoders: A Comprehensive Review
- Audio-visual Speech Enhancement Using Conditional Variational Auto-Encoders
- Semi-supervised multichannel speech enhancement with variational autoencoders and non-negative matrix factorization
- Unsupervised Speech Enhancement Based on Multichannel NMF-Informed Beamforming for Noise-Robust Automatic Speech Recognition
- Variational Autoencoder for Speech Enhancement with a Noise-Aware Encoder
- A variance modeling framework based on variational autoencoders for speech enhancement
- Speech enhancement with variational autoencoders and alpha-stable distributions
- The Ethical Implications of Generative Audio Models: A Systematic Literature Review
- Semi-blind source separation with multichannel variational autoencoder
- Mixture of Inference Networks for VAE-based Audio-visual Speech Enhancement
- End-to-End Multi-Task Denoising for joint SDR and PESQ Optimization
- Investigating the Design Space of Diffusion Models for Speech Enhancement
- A Deep Generative Model of Speech Complex Spectrograms
- Guided Variational Autoencoder for Speech Enhancement With a Supervised Classifier
- Deep Speech Enhancement for Reverberated and Noisy Signals using Wide Residual Networks
- Disentanglement Learning for Variational Autoencoders Applied to Audio-Visual Speech Enhancement
- Learning and controlling the source-filter representation of speech with a variational autoencoder
- Estimation with Low-Rank Time-Frequency Synthesis Models
- Integrating Uncertainty into Neural Network-based Speech Enhancement
- Integrating Statistical Uncertainty into Neural Network-Based Speech Enhancement
- RemixIT: Continual self-training of speech enhancement models via bootstrapped remixing
- Learning robust speech representation with an articulatory-regularized variational autoencoder
- Fast Multichannel Source Separation Based on Jointly Diagonalizable Spatial Covariance Matrices
- A Speech Enhancement Algorithm based on Non-negative Hidden Markov Model and Kullback-Leibler Divergence
- Probabilistic Modelling of Signal Mixtures with Differentiable Dictionaries
- FastMVAE2: On improving and accelerating the fast variational autoencoder-based source separation algorithm for determined mixtures
- Fast MVAE: Joint separation and classification of mixed sources based on multichannel variational autoencoder with auxiliary classifier
- A Benchmark of Dynamical Variational Autoencoders applied to Speech Spectrogram Modeling
- DenDrift: A Drift-Aware Algorithm for Host Profiling
- Can We Trust Deep Speech Prior?