Conditional Generative Adversarial Networks for Speech Enhancement and Noise-Robust Speaker Verification
arXiv:1709.01703 · doi:10.21437/Interspeech.2017-1620
Abstract
Improving speech system performance in noisy environments remains a challenging task, and speech enhancement (SE) is one of the effective techniques to solve the problem. Motivated by the promising results of generative adversarial networks (GANs) in a variety of image processing tasks, we explore the potential of conditional GANs (cGANs) for SE, and in particular, we make use of the image processing framework proposed by Isola et al. [1] to learn a mapping from the spectrogram of noisy speech to an enhanced counterpart. The SE cGAN consists of two networks, trained in an adversarial manner: a generator that tries to enhance the input noisy spectrogram, and a discriminator that tries to distinguish between enhanced spectrograms provided by the generator and clean ones from the database using the noisy spectrogram as a condition. We evaluate the performance of the cGAN method in terms of perceptual evaluation of speech quality (PESQ), short-time objective intelligibility (STOI), and equal error rate (EER) of speaker verification (an example application). Experimental results show that the cGAN method overall outperforms the classical short-time spectral amplitude minimum mean square error (STSA-MMSE) SE algorithm, and is comparable to a deep neural network-based SE approach (DNN-SE).
INTERSPEECH 2017 August 20-24, 2017, Stockholm, Sweden
References in corpus (7)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Conditional Generative Adversarial Nets
- Striving for Simplicity: The All Convolutional Net
- NIPS 2016 Tutorial: Generative Adversarial Networks
- C-RNN-GAN: Continuous recurrent neural networks with adversarial training
- Image De-raining Using a Conditional Generative Adversarial Network
- Voice Conversion using Convolutional Neural Networks
Cited by in corpus (53)
- Deep Learning for Audio Signal Processing
- CMGAN: Conformer-Based Metric-GAN for Monaural Speech Enhancement
- Characterizing Audio Adversarial Examples Using Temporal Dependency
- Probabilistic Forecasting of Sensory Data with Generative Adversarial Networks - ForGAN
- Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends
- Characterizing Speech Adversarial Examples Using Self-Attention U-Net Enhancement
- Deep Learning for Environmentally Robust Speech Recognition: An Overview of Recent Developments
- Supervised Speech Separation Based on Deep Learning: An Overview
- Cross-lingual Text-independent Speaker Verification using Unsupervised Adversarial Discriminative Domain Adaptation
- Time-Domain Multi-modal Bone/air Conducted Speech Enhancement
- Deep-Learning-Based Audio-Visual Speech Enhancement in Presence of Lombard Effect
- UNetGAN: A Robust Speech Enhancement Approach in Time Domain for Extremely Low Signal-to-noise Ratio Condition
- Investigating Generative Adversarial Networks based Speech Dereverberation for Robust Speech Recognition
- Time-Contrastive Learning Based Deep Bottleneck Features for Text-Dependent Speaker Verification
- Stochastic Restoration of Heavily Compressed Musical Audio using Generative Adversarial Networks
- SkipConvGAN: Monaural Speech Dereverberation using Generative Adversarial Networks via Complex Time-Frequency Masking
- SNR-Based Features and Diverse Training Data for Robust DNN-Based Speech Enhancement
- End-to-End Waveform Utterance Enhancement for Direct Evaluation Metrics Optimization by Fully Convolutional Neural Networks
- Investigating the Design Space of Diffusion Models for Speech Enhancement
- A Study on Speech Enhancement Based on Diffusion Probabilistic Model
- Libri-Adapt: A New Speech Dataset for Unsupervised Domain Adaptation
- Noise Adaptive Speech Enhancement using Domain Adversarial Training
- Deep Speech Enhancement for Reverberated and Noisy Signals using Wide Residual Networks
- DrumGAN: Synthesis of Drum Sounds With Timbral Feature Conditioning Using Generative Adversarial Networks
- Speech Enhancement based on Denoising Autoencoder with Multi-branched Encoders
- Speaker Recognition Based on Deep Learning: An Overview
- High-quality Speech Synthesis Using Super-resolution Mel-Spectrogram
- iSEGAN: Improved Speech Enhancement Generative Adversarial Networks
- INTERSPEECH 2021 ConferencingSpeech Challenge: Towards Far-field Multi-Channel Speech Enhancement for Video Conferencing
- Improved Lite Audio-Visual Speech Enhancement
- Distributed Microphone Speech Enhancement based on Deep Learning
- Dynamic Attention Based Generative Adversarial Network with Phase Post-Processing for Speech Enhancement
- rVAD: An Unsupervised Segment-Based Robust Voice Activity Detection Method
- Improving Speech Enhancement Performance by Leveraging Contextual Broad Phonetic Class Information
- Investigating Cross-Domain Losses for Speech Enhancement
- Adversarial Example Detection by Classification for Deep Speech Recognition
- Late reverberation suppression using U-nets
- SADDEL: Joint Speech Separation and Denoising Model based on Multitask Learning
- Adversarial Joint Training with Self-Attention Mechanism for Robust End-to-End Speech Recognition
- DarkGAN: Exploiting Knowledge Distillation for Comprehensible Audio Synthesis with GANs
- Adversarial Data Augmentation for Disordered Speech Recognition
- Reducing over-smoothness in speech synthesis using Generative Adversarial Networks
- MTGAN: Speaker Verification through Multitasking Triplet Generative Adversarial Networks
- Generative Speech Enhancement Based on Cloned Networks
- Data augmentation enhanced speaker enrollment for text-dependent speaker verification
- Normalized Features for Improving the Generalization of DNN Based Speech Enhancement
- Visual Speech Enhancement Without A Real Visual Stream
- Incorporating Symbolic Sequential Modeling for Speech Enhancement
- BiNet: Degraded-Manuscript Binarization in Diverse Document Textures and Layouts using Deep Encoder-Decoder Networks
- Data Generation Using Pass-phrase-dependent Deep Auto-encoders for Text-Dependent Speaker Verification
- Speaker Representation Learning using Global Context Guided Channel and Time-Frequency Transformations
- Data augmentation using generative networks to identify dementia
- Coarse-to-fine Optimization for Speech Enhancement