MetricGAN: Generative Adversarial Networks based Black-box Metric Scores Optimization for Speech Enhancement
arXiv:1905.04874
Abstract
Adversarial loss in a conditional generative adversarial network (GAN) is not designed to directly optimize evaluation metrics of a target task, and thus, may not always guide the generator in a GAN to generate data with improved metric scores. To overcome this issue, we propose a novel MetricGAN approach with an aim to optimize the generator with respect to one or multiple evaluation metrics. Moreover, based on MetricGAN, the metric scores of the generated data can also be arbitrarily specified by users. We tested the proposed MetricGAN on a speech enhancement task, which is particularly suitable to verify the proposed approach because there are multiple metrics measuring different aspects of speech signals. Moreover, these metrics are generally complex and could not be fully optimized by Lp or conventional adversarial losses.
Accepted by Thirty-sixth International Conference on Machine Learning (ICML) 2019
Cited by in corpus (18)
- ClearBuds: Wireless Binaural Earbuds for Learning-Based Speech Enhancement
- Stochastic Restoration of Heavily Compressed Musical Audio using Generative Adversarial Networks
- NORESQA: A Framework for Speech Quality Assessment using Non-Matching References
- HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial Networks
- CycleGAN-based Non-parallel Speech Enhancement with an Adaptive Attention-in-attention Mechanism
- STOI-Net: A Deep Learning based Non-Intrusive Speech Intelligibility Assessment Model
- Glance and Gaze: A Collaborative Learning Framework for Single-channel Speech Enhancement
- TSTNN: Two-stage Transformer based Neural Network for Speech Enhancement in the Time Domain
- Speech Enhancement using Self-Adaptation and Multi-Head Self-Attention
- Real-time speech enhancement using equilibriated RNN
- A Two-stage Complex Network using Cycle-consistent Generative Adversarial Networks for Speech Enhancement
- Dynamic Attention Based Generative Adversarial Network with Phase Post-Processing for Speech Enhancement
- Speech Enhancement Using Multi-Stage Self-Attentive Temporal Convolutional Networks
- Masks Fusion with Multi-Target Learning For Speech Enhancement
- iMetricGAN: Intelligibility Enhancement for Speech-in-Noise using Generative Adversarial Network-based Metric Learning
- SERIL: Noise Adaptive Speech Enhancement using Regularization-based Incremental Learning
- Stable Training of DNN for Speech Enhancement based on Perceptually-Motivated Black-Box Cost Function
- Invertible DNN-based nonlinear time-frequency transform for speech enhancement