StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation
arXiv:2212.11851 · doi:10.1109/TASLP.2023.3294692
Abstract
Diffusion models have shown a great ability at bridging the performance gap between predictive and generative approaches for speech enhancement. We have shown that they may even outperform their predictive counterparts for non-additive corruption types or when they are evaluated on mismatched conditions. However, diffusion models suffer from a high computational burden, mainly as they require to run a neural network for each reverse diffusion step, whereas predictive approaches only require one pass. As diffusion models are generative approaches they may also produce vocalizing and breathing artifacts in adverse conditions. In comparison, in such difficult scenarios, predictive models typically do not produce such artifacts but tend to distort the target speech instead, thereby degrading the speech quality. In this work, we present a stochastic regeneration approach where an estimate given by a predictive model is provided as a guide for further diffusion. We show that the proposed approach uses the predictive model to remove the vocalizing and breathing artifacts while producing very high quality samples thanks to the diffusion model, even in adverse conditions. We further show that this approach enables to use lighter sampling schemes with fewer diffusion steps without sacrificing quality, thus lifting the computational burden by an order of magnitude. Source code and audio examples are available online (https://uhh.de/inf-sp-storm).
Published in IEEE/ACM Transactions on Audio, Speech and Language Processing, 2023
References in corpus (8)
- Diffusion Models Beat GANs on Image Synthesis
- Score-Based Generative Modeling through Stochastic Differential Equations
- Speech Enhancement and Dereverberation with Diffusion-based Generative Models
- A variance modeling framework based on variational autoencoders for speech enhancement
- NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates
- Inversion by Direct Iteration: An Alternative to Denoising Diffusion for Image Restoration
- BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis
- VoiceFixer: Toward General Speech Restoration with Neural Vocoder
Cited by in corpus (7)
- Investigating the Design Space of Diffusion Models for Speech Enhancement
- DriftRec: Adapting diffusion models to blind JPEG restoration
- AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
- Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
- Crowdsourced Multilingual Speech Intelligibility Testing
- GALD-SE: Guided Anisotropic Lightweight Diffusion for Efficient Speech Enhancement
- Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement