SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
arXiv:1612.07837
Abstract
In this paper we propose a novel model for unconditional audio generation based on generating one audio sample at a time. We show that our model, which profits from combining memory-less modules, namely autoregressive multilayer perceptrons, and stateful recurrent neural networks in a hierarchical structure is able to capture underlying sources of variations in the temporal sequences over very long time spans, on three datasets of different nature. Human evaluation on the generated samples indicate that our model is preferred over competing models. We also show how each component of the model contributes to the exhibited performance.
Published as a conference paper at ICLR 2017
Cited by in corpus (40)
- Deep Learning for Audio Signal Processing
- Implicit Neural Representations with Periodic Activation Functions
- GANSynth: Adversarial Neural Audio Synthesis
- A Survey on Neural Speech Synthesis
- Jukebox: A Generative Model for Music
- High Fidelity Speech Synthesis with Adversarial Networks
- A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions
- DDSP: Differentiable Digital Signal Processing
- Waveform Modeling and Generation Using Hierarchical Recurrent Neural Networks for Speech Bandwidth Extension
- Codified audio language modeling learns useful representations for music information retrieval
- Taming Visually Guided Sound Generation
- A general-purpose deep learning approach to model time-varying audio effects
- Multi-Instrumentalist Net: Unsupervised Generation of Music from Body Movements
- Adversarial Generation of Time-Frequency Features with application in audio synthesis
- Cross-speaker Emotion Transfer Based on Speaker Condition Layer Normalization and Semi-Supervised Training in Text-To-Speech
- SingSong: Generating musical accompaniments from singing
- RAVE: A variational autoencoder for fast and high-quality neural audio synthesis
- Catch-A-Waveform: Learning to Generate Audio from a Single Short Example
- Non-parallel Voice Conversion System with WaveNet Vocoder and Collapsed Speech Suppression
- From Speaker Verification to Multispeaker Speech Synthesis, Deep Transfer with Feedback Constraint
- Diet deep generative audio models with structured lottery
- Vertical-Horizontal Structured Attention for Generating Music with Chords
- Knowledge-and-Data-Driven Amplitude Spectrum Prediction for Hierarchical Neural Vocoders
- Multi-scale Transformer Language Models
- Vision-Infused Deep Audio Inpainting
- Glow-WaveGAN: Learning Speech Representations from GAN-based Variational Auto-Encoder For High Fidelity Flow-based Speech Synthesis
- Source-Filter-Based Generative Adversarial Neural Vocoder for High Fidelity Speech Synthesis
- On-device neural speech synthesis
- The AS-NU System for the M2VoC Challenge
- I'm Sorry for Your Loss: Spectrally-Based Audio Distances Are Bad at Pitch
- Quasi-Periodic Parallel WaveGAN Vocoder: A Non-autoregressive Pitch-dependent Dilated Convolution Model for Parametric Speech Generation
- Conditional Sound Generation Using Neural Discrete Time-Frequency Representation Learning
- Distilling the Knowledge from Conditional Normalizing Flows
- NeuralDPS: Neural Deterministic Plus Stochastic Model with Multiband Excitation for Noise-Controllable Waveform Generation
- Improve GAN-based Neural Vocoder using Pointwise Relativistic LeastSquare GAN
- Transferring neural speech waveform synthesizers to musical instrument sounds generation
- End-To-End Dilated Variational Autoencoder with Bottleneck Discriminative Loss for Sound Morphing -- A Preliminary Study
- SchrödingeRNN: Generative Modeling of Raw Audio as a Continuously Observed Quantum State
- Text-to-speech for the hearing impaired
- Relational Data Selection for Data Augmentation of Speaker-dependent Multi-band MelGAN Vocoder