Diffusion-Based Audio Inpainting
arXiv:2305.15266 · doi:10.17743/jaes.2022.0129
Abstract
Audio inpainting aims to reconstruct missing segments in corrupted recordings. Most of existing methods produce plausible reconstructions when the gap lengths are short, but struggle to reconstruct gaps larger than about 100 ms. This paper explores recent advancements in deep learning and, particularly, diffusion models, for the task of audio inpainting. The proposed method uses an unconditionally trained generative model, which can be conditioned in a zero-shot fashion for audio inpainting, and is able to regenerate gaps of any size. An improved deep neural network architecture based on the constant-Q transform, which allows the model to exploit pitch-equivariant symmetries in audio, is also presented. The performance of the proposed algorithm is evaluated through objective and subjective metrics for the task of reconstructing short to mid-sized gaps, up to 300 ms. The results of a formal listening test show that the proposed method delivers comparable performance against the compared baselines for short gaps, such as 50 ms, while retaining a good audio quality and outperforming the baselines for wider gaps that are up to 300 ms long. The method presented in this paper can be applied to restoring sound recordings that suffer from severe local disturbances or dropouts, which must be reconstructed.
Submitted for publication to the Journal of Audio Engineering Society on January 30th, 2023
References in corpus (23)
- Denoising Diffusion Probabilistic Models
- Diffusion Models Beat GANs on Image Synthesis
- Score-Based Generative Modeling through Stochastic Differential Equations
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains
- Elucidating the Design Space of Diffusion-Based Generative Models
- Speech Enhancement and Dereverberation with Diffusion-based Generative Models
- Denoising Diffusion Restoration Models
- FiLM: Visual Reasoning with a General Conditioning Layer
- Diffusion Posterior Sampling for General Noisy Inverse Problems
- Learning Features of Music from Scratch
- DiffWave: A Versatile Diffusion Model for Audio Synthesis
- Improving Diffusion Models for Inverse Problems using Manifold Constraints
- A framework for invertible, real-time constant-Q transforms
- Zero-Shot Image Restoration Using Denoising Diffusion Null-Space Model
- How Much Position Information Do Convolutional Neural Networks Encode?
- RePaint: Inpainting using Denoising Diffusion Probabilistic Models
- GACELA -- A generative adversarial context encoder for long audio inpainting
- Introducing SPAIN (SParse Audio INpainter)
- Audio inpainting with generative adversarial network
- Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier Guidance
- AUDIT: Audio Editing by Following Instructions with Latent Diffusion Models
- SpeechPainter: Text-conditioned Speech Inpainting
- Rethinking complex-valued deep neural networks for monaural speech enhancement