GACELA -- A generative adversarial context encoder for long audio inpainting
arXiv:2005.05032 · doi:10.1109/JSTSP.2020.3037506
Abstract
We introduce GACELA, a generative adversarial network (GAN) designed to restore missing musical audio data with a duration ranging between hundreds of milliseconds to a few seconds, i.e., to perform long-gap audio inpainting. While previous work either addressed shorter gaps or relied on exemplars by copying available information from other signal parts, GACELA addresses the inpainting of long gaps in two aspects. First, it considers various time scales of audio information by relying on five parallel discriminators with increasing resolution of receptive fields. Second, it is conditioned not only on the available information surrounding the gap, i.e., the context, but also on the latent variable of the conditional GAN. This addresses the inherent multi-modality of audio inpainting at such long gaps and provides the option of user-defined inpainting. GACELA was tested in listening tests on music signals of varying complexity and gap durations ranging from 375~ms to 1500~ms. While our subjects were often able to detect the inpaintings, the severity of the artifacts decreased from unacceptable to mildly disturbing. GACELA represents a framework capable to integrate future improvements such as processing of more auditory-related features or more explicit musical features.
References in corpus (17)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- WaveNet: A Generative Model for Raw Audio
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- Modeling Temporal Dependencies in High-Dimensional Sequences: Application to Polyphonic Music Generation and Transcription
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
- GANSynth: Adversarial Neural Audio Synthesis
- Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset
- A Functional Taxonomy of Music Generation Systems
- The challenge of realistic music generation: modelling raw audio at scale
- MelNet: A Generative Model for Audio in the Frequency Domain
- High Fidelity Speech Synthesis with Adversarial Networks
- Deep speech inpainting of time-frequency masks
- Introducing SPAIN (SParse Audio INpainter)
- Deep Long Audio Inpainting
- Audio Inpainting: Revisited and Reweighted
- Audio inpainting of music by means of neural networks
- Conditioning Deep Generative Raw Audio Models for Structured Automatic Music
Cited by in corpus (10)
- A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions
- A survey and an extensive evaluation of popular audio declipping methods
- Time-Frequency Phase Retrieval for Audio -- The Effect of Transform Parameters
- Diffusion-Based Audio Inpainting
- Catch-A-Waveform: Learning to Generate Audio from a Single Short Example
- Multiple Hankel matrix rank minimization for audio inpainting
- Algorithms for audio inpainting based on probabilistic nonnegative matrix factorization
- Tweaking autoregressive methods for inpainting of gaps in audio signals
- Phase-Based Signal Representations for Scattering
- Janssen 2.0: Audio Inpainting in the Time-frequency Domain