A unified convolutional beamformer for simultaneous denoising and dereverberation
arXiv:1812.08400 · doi:10.1109/LSP.2019.2911179
Abstract
This paper proposes a method for estimating a convolutional beamformer that can perform denoising and dereverberation simultaneously in an optimal way. The application of dereverberation based on a weighted prediction error (WPE) method followed by denoising based on a minimum variance distortionless response (MVDR) beamformer has conventionally been considered a promising approach, however, the optimality of this approach cannot be guaranteed. To realize the optimal integration of denoising and dereverberation, we present a method that unifies the WPE dereverberation method and a variant of the MVDR beamformer, namely a minimum power distortionless response (MPDR) beamformer, into a single convolutional beamformer, and we optimize it based on a single unified optimization criterion. The proposed beamformer is referred to as a Weighted Power minimization Distortionless response (WPD) beamformer. Experiments show that the proposed method substantially improves the speech enhancement performance in terms of both objective speech enhancement measures and automatic speech recognition (ASR) performance.
Published in IEEE Signal Processing Letters
Cited by in corpus (7)
- Twenty-Five Years of Advances in Beamforming: From Convex and Nonconvex Optimization to Learning Techniques
- Jointly optimal denoising, dereverberation, and source separation
- Integrated sidelobe cancellation and linear prediction Kalman filter for joint multi-microphone speech dereverberation, interfering speech cancellation, and noise reduction
- Exploring the time-domain deep attractor network with two-stream architectures in a reverberant environment
- ESPnet-se: end-to-end speech enhancement and separation toolkit designed for asr integration
- Adaptive Dereverberation, Noise and Interferer Reduction Using Sparse Weighted Linearly Constrained Minimum Power Beamforming
- End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend