Learning with Learned Loss Function: Speech Enhancement with Quality-Net to Improve Perceptual Evaluation of Speech Quality
arXiv:1905.01898 · doi:10.1109/LSP.2019.2953810
Abstract
Utilizing a human-perception-related objective function to train a speech enhancement model has become a popular topic recently. The main reason is that the conventional mean squared error (MSE) loss cannot represent auditory perception well. One of the typical hu-man-perception-related metrics, which is the perceptual evaluation of speech quality (PESQ), has been proven to provide a high correlation to the quality scores rated by humans. Owing to its complex and non-differentiable properties, however, the PESQ function may not be used to optimize speech enhancement models directly. In this study, we propose optimizing the enhancement model with an approximated PESQ function, which is differentiable and learned from the training data. The experimental results show that the learned surrogate function can guide the enhancement model to further boost the PESQ score (in-crease of 0.18 points compared to the results trained with MSE loss) and maintain the speech intelligibility.
Accepted by IEEE Signal Processing Letters (SPL)
References in corpus (3)
Cited by in corpus (16)
- MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement
- Audio Description from Image by Modal Translation Network
- HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial Networks
- Boosting Objective Scores of a Speech Enhancement Model by MetricGAN Post-processing
- Improving Perceptual Quality by Phone-Fortified Perceptual Loss using Wasserstein Distance for Speech Enhancement
- Learning to Maximize Speech Quality Directly Using MOS Prediction for Neural Text-to-Speech
- Controlling the Remixing of Separated Dialogue with a Non-Intrusive Quality Estimate
- Capacity-Net-Based RIS Precoding Design without Channel Estimation for mmWave MIMO System
- A Differentiable Perceptual Audio Metric Learned from Just Noticeable Differences
- Deep Noise Suppression With Non-Intrusive PESQNet Supervision Enabling the Use of Real Training Data
- Frequency Gating: Improved Convolutional Neural Networks for Speech Enhancement in the Time-Frequency Domain
- dCoNNear: An Artifact-Free Neural Network Architecture for Closed-loop Audio Signal Processing
- Masks Fusion with Multi-Target Learning For Speech Enhancement
- InSE-NET: A Perceptually Coded Audio Quality Model based on CNN
- Phase Aware Speech Enhancement using Realisation of Complex-valued LSTM
- A Deep Learning Loss Function based on Auditory Power Compression for Speech Enhancement