Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
arXiv:2210.15793 · doi:10.1109/ICASSP49357.2023.10095103
Abstract
Recently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given low-resolution speech utterances. This is commonly achieved by conditioning the network of noise predictor with low-resolution audio. In this paper, we propose a novel sampling algorithm that communicates the information of the low-resolution audio via the reverse sampling process of DMs. The proposed method can be a drop-in replacement for the vanilla sampling process and can significantly improve the performance of the existing works. Moreover, by coupling the proposed sampling method with an unconditional DM, i.e., a DM with no auxiliary inputs to its noise predictor, we can generalize it to a wide range of SR setups. We also attain state-of-the-art results on the VCTK Multi-Speaker benchmark with this novel formulation.
Published at ICASSP 2023
References in corpus (8)
- Score-Based Generative Modeling through Stochastic Differential Equations
- Variational Diffusion Models
- Denoising Diffusion Restoration Models
- Solving Inverse Problems in Medical Imaging with Score-Based Generative Models
- Improving Diffusion Models for Inverse Problems using Manifold Constraints
- NU-Wave: A Diffusion Probabilistic Model for Neural Audio Upsampling
- NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates
- A Study on Speech Enhancement Based on Diffusion Probabilistic Model