CMGAN: Conformer-Based Metric-GAN for Monaural Speech Enhancement
arXiv:2209.11112 · doi:10.1109/TASLP.2024.3393718
Abstract
In this work, we further develop the conformer-based metric generative adversarial network (CMGAN) model for speech enhancement (SE) in the time-frequency (TF) domain. This paper builds on our previous work but takes a more in-depth look by conducting extensive ablation studies on model inputs and architectural design choices. We rigorously tested the generalization ability of the model to unseen noise types and distortions. We have fortified our claims through DNS-MOS measurements and listening tests. Rather than focusing exclusively on the speech denoising task, we extend this work to address the dereverberation and super-resolution tasks. This necessitated exploring various architectural changes, specifically metric discriminator scores and masking techniques. It is essential to highlight that this is among the earliest works that attempted complex TF-domain super-resolution. Our findings show that CMGAN outperforms existing state-of-the-art methods in the three major speech enhancement tasks: denoising, dereverberation, and super-resolution. For example, in the denoising task using the Voice Bank+DEMAND dataset, CMGAN notably exceeded the performance of prior models, attaining a PESQ score of 3.41 and an SSNR of 11.10 dB. Audio samples and CMGAN implementations are available online.
17 pages, 11 figures, and 6 tables. arXiv admin note: text overlap with arXiv:2203.15149
References in corpus (9)
- Deep Learning for Audio Signal Processing
- CMGAN: Conformer-based Metric GAN for Speech Enhancement
- On The Compensation Between Magnitude and Phase in Speech Separation
- Universal Speech Enhancement with Score-based Diffusion
- Neural Vocoder is All You Need for Speech Super-resolution
- SkipConvNet: Skip Convolutional Neural Network for Speech Dereverberation using Optimally Smoothed Spectral Mapping
- Multi-Domain Processing via Hybrid Denoising Networks for Speech Enhancement
- SkipConvGAN: Monaural Speech Dereverberation using Generative Adversarial Networks via Complex Time-Frequency Masking
- Deep Speech Enhancement for Reverberated and Noisy Signals using Wide Residual Networks
Cited by in corpus (4)
- TRNet: Two-level Refinement Network leveraging Speech Enhancement for Noise Robust Speech Emotion Recognition
- TF-CorrNet: Leveraging Spatial Correlation for Continuous Speech Separation
- Study of Lightweight Transformer Architectures for Single-Channel Speech Enhancement
- TalkLess: Blending Extractive and Abstractive Speech Summarization for Editing Speech to Preserve Content and Style