Speech Enhancement and Dereverberation with Diffusion-based Generative Models
arXiv:2208.05830 · doi:10.1109/TASLP.2023.3285241
Abstract
In this work, we build upon our previous publication and use diffusion-based generative models for speech enhancement. We present a detailed overview of the diffusion process that is based on a stochastic differential equation and delve into an extensive theoretical examination of its implications. Opposed to usual conditional generation tasks, we do not start the reverse process from pure Gaussian noise but from a mixture of noisy speech and Gaussian noise. This matches our forward process which moves from clean speech to noisy speech by including a drift term. We show that this procedure enables using only 30 diffusion steps to generate high-quality clean speech estimates. By adapting the network architecture, we are able to significantly improve the speech enhancement performance, indicating that the network, rather than the formalism, was the main limitation of our original approach. In an extensive cross-dataset evaluation, we show that the improved method can compete with recent discriminative models and achieves better generalization when evaluating on a different corpus than used for training. We complement the results with an instrumental evaluation using real-world noisy recordings and a listening experiment, in which our proposed method is rated best. Examining different sampler configurations for solving the reverse process allows us to balance the performance and computational speed of the proposed method. Moreover, we show that the proposed method is also suitable for dereverberation and thus not limited to additive background noise removal. Code and audio examples are available online, see https://github.com/sp-uhh/sgmse.
Proofread version
References in corpus (9)
- Generative Adversarial Networks
- Diffusion Models Beat GANs on Image Synthesis
- Elucidating the Design Space of Diffusion-Based Generative Models
- CMGAN: Conformer-based Metric GAN for Speech Enhancement
- Universal Speech Enhancement with Score-based Diffusion
- A variance modeling framework based on variational autoencoders for speech enhancement
- HiFi++: a Unified Framework for Bandwidth Extension and Speech Enhancement
- A Variational Perspective on Diffusion-Based Generative Models and Score Matching
- Phase-Aware Deep Speech Enhancement: It's All About The Frame Length
Cited by in corpus (23)
- StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation
- SpectralDiff: A Generative Framework for Hyperspectral Image Classification with Diffusion Models
- The Ethical Implications of Generative Audio Models: A Systematic Literature Review
- Investigating the Design Space of Diffusion Models for Speech Enhancement
- Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge
- DriftRec: Adapting diffusion models to blind JPEG restoration
- A Multi-dimensional Deep Structured State Space Approach to Speech Enhancement Using Small-footprint Models
- DiffPhase: Generative Diffusion-based STFT Phase Retrieval
- AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
- Diffiner: A Versatile Diffusion-based Generative Refiner for Speech Enhancement
- Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
- Diffusion-Based Audio Inpainting
- Crowdsourced Multilingual Speech Intelligibility Testing
- The Effect of Training Dataset Size on Discriminative and Diffusion-Based Speech Enhancement Systems
- Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler
- SEFGAN: Harvesting the Power of Normalizing Flows and GANs for Efficient High-Quality Speech Enhancement
- An empirical study on speech restoration guided by self supervised speech representation
- A Dual-Branch Parallel Network for Speech Enhancement and Restoration
- Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
- GALD-SE: Guided Anisotropic Lightweight Diffusion for Efficient Speech Enhancement
- Room-acoustic simulations as an alternative to measurements for audio-algorithm evaluation
- From Physics to Representation: Audio Learning with Synthetic Pre-training via Procedural Generation
- Towards Reliable Objective Evaluation Metrics for Generative Singing Voice Separation Models