MetricGAN: Generative Adversarial Networks based Black-box Metric Scores Optimization for Speech Enhancement
arXiv:1905.04874
Abstract
Adversarial loss in a conditional generative adversarial network (GAN) is not designed to directly optimize evaluation metrics of a target task, and thus, may not always guide the generator in a GAN to generate data with improved metric scores. To overcome this issue, we propose a novel MetricGAN approach with an aim to optimize the generator with respect to one or multiple evaluation metrics. Moreover, based on MetricGAN, the metric scores of the generated data can also be arbitrarily specified by users. We tested the proposed MetricGAN on a speech enhancement task, which is particularly suitable to verify the proposed approach because there are multiple metrics measuring different aspects of speech signals. Moreover, these metrics are generally complex and could not be fully optimized by Lp or conventional adversarial losses.
Accepted by Thirty-sixth International Conference on Machine Learning (ICML) 2019
Cited by in corpus (31)
- Self-attending RNN for Speech Enhancement to Improve Cross-corpus Generalization
- ClearBuds: Wireless Binaural Earbuds for Learning-Based Speech Enhancement
- Stochastic Restoration of Heavily Compressed Musical Audio using Generative Adversarial Networks
- MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement
- NORESQA: A Framework for Speech Quality Assessment using Non-Matching References
- Listening to Sounds of Silence for Speech Denoising
- HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial Networks
- A Study on Speech Enhancement Based on Diffusion Probabilistic Model
- Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge
- CycleGAN-based Non-parallel Speech Enhancement with an Adaptive Attention-in-attention Mechanism
- STOI-Net: A Deep Learning based Non-Intrusive Speech Intelligibility Assessment Model
- Glance and Gaze: A Collaborative Learning Framework for Single-channel Speech Enhancement
- Improving Perceptual Quality by Phone-Fortified Perceptual Loss using Wasserstein Distance for Speech Enhancement
- Dynamic Noise Embedding: Noise Aware Training and Adaptation for Speech Enhancement
- TSTNN: Two-stage Transformer based Neural Network for Speech Enhancement in the Time Domain
- Controlling the Remixing of Separated Dialogue with a Non-Intrusive Quality Estimate
- Speech Enhancement using Self-Adaptation and Multi-Head Self-Attention
- A Two-stage Complex Network using Cycle-consistent Generative Adversarial Networks for Speech Enhancement
- Real-time speech enhancement using equilibriated RNN
- Dynamic Attention Based Generative Adversarial Network with Phase Post-Processing for Speech Enhancement
- Speech Enhancement Using Multi-Stage Self-Attentive Temporal Convolutional Networks
- Time-domain Speech Enhancement with Generative Adversarial Learning
- Deep Learning-based Non-Intrusive Multi-Objective Speech Assessment Model with Cross-Domain Features
- iMetricGAN: Intelligibility Enhancement for Speech-in-Noise using Generative Adversarial Network-based Metric Learning
- Masks Fusion with Multi-Target Learning For Speech Enhancement
- Invertible DNN-based nonlinear time-frequency transform for speech enhancement
- SERIL: Noise Adaptive Speech Enhancement using Regularization-based Incremental Learning
- Stable Training of DNN for Speech Enhancement based on Perceptually-Motivated Black-Box Cost Function
- Speech Enhancement Based on Cyclegan with Noise-informed Training
- Multi-Metric Optimization using Generative Adversarial Networks for Near-End Speech Intelligibility Enhancement
- Speech Enhancement using Separable Polling Attention and Global Layer Normalization followed with PReLU