Phase-aware Speech Enhancement with Deep Complex U-Net
arXiv:1903.03107
Abstract
Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of clean speech. To improve speech enhancement performance, we tackle the phase estimation problem in three ways. First, we propose Deep Complex U-Net, an advanced U-Net structured model incorporating well-defined complex-valued building blocks to deal with complex-valued spectrograms. Second, we propose a polar coordinate-wise complex-valued masking method to reflect the distribution of complex ideal ratio masks. Third, we define a novel loss function, weighted source-to-distortion ratio (wSDR) loss, which is designed to directly correlate with a quantitative evaluation measure. Our model was evaluated on a mixture of the Voice Bank corpus and DEMAND database, which has been widely used by many deep learning models for speech enhancement. Ablation experiments were conducted on the mixed dataset showing that all three proposed approaches are empirically valid. Experimental results show that the proposed method achieves state-of-the-art performance in all metrics, outperforming previous approaches by a large margin.
Significant error was found in data processing step, therefore will be retracted from International Conference on Learning Representations (ICLR) 2019. It is not recommended to read current version
Cited by in corpus (10)
- End-to-End Multi-Channel Speech Separation
- DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement
- Temporal-Spatial Neural Filter: Direction Informed End-to-End Multi-channel Target Speech Separation
- TSTNN: Two-stage Transformer based Neural Network for Speech Enhancement in the Time Domain
- Phase-aware Single-stage Speech Denoising and Dereverberation with U-Net
- RHR-Net: A Residual Hourglass Recurrent Neural Network for Speech Enhancement
- Data Efficient Voice Cloning from Noisy Samples with Domain Adversarial Training
- Improved Speech Separation with Time-and-Frequency Cross-domain Joint Embedding and Clustering
- Real-time Streaming Wave-U-Net with Temporal Convolutions for Multichannel Speech Enhancement
- End-to-end speech enhancement based on discrete cosine transform