Neural Synthesis of Footsteps Sound Effects with Generative Adversarial Networks
arXiv:2110.09605
Abstract
Footsteps are among the most ubiquitous sound effects in multimedia applications. There is substantial research into understanding the acoustic features and developing synthesis models for footstep sound effects. In this paper, we present a first attempt at adopting neural synthesis for this task. We implemented two GAN-based architectures and compared the results with real recordings as well as six traditional sound synthesis methods. Our architectures reached realism scores as high as recorded samples, showing encouraging results for the task at hand.
References in corpus (7)
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- Universal MelGAN: A Robust Neural Vocoder for High-Fidelity Waveform Generation in Multiple Domains
- DrumGAN: Synthesis of Drum Sounds With Timbral Feature Conditioning Using Generative Adversarial Networks
- VocGAN: A High-Fidelity Real-time Vocoder with a Hierarchically-nested Adversarial Network
- GAN Vocoder: Multi-Resolution Discriminator Is All You Need
- Improved parallel WaveGAN vocoder with perceptually weighted spectrogram loss