A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions
arXiv:2011.06801
Abstract
The utilization of deep learning techniques in generating various contents (such as image, text, etc.) has become a trend. Especially music, the topic of this paper, has attracted widespread attention of countless researchers.The whole process of producing music can be divided into three stages, corresponding to the three levels of music generation: score generation produces scores, performance generation adds performance characteristics to the scores, and audio generation converts scores with performance characteristics into audio by assigning timbre or generates music in audio format directly. Previous surveys have explored the network models employed in the field of automatic music generation. However, the development history, the model evolution, as well as the pros and cons of same music generation task have not been clearly illustrated. This paper attempts to provide an overview of various composition tasks under different music generation levels, covering most of the currently popular music generation tasks using deep learning. In addition, we summarize the datasets suitable for diverse tasks, discuss the music representations, the evaluation methods as well as the challenges under different levels, and finally point out several future directions.
96 pages,this is a draft
References in corpus (35)
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Conditional Generative Adversarial Nets
- From Frequency to Meaning: Vector Space Models of Semantics
- Modeling Temporal Dependencies in High-Dimensional Sequences: Application to Polyphonic Music Generation and Transcription
- C-RNN-GAN: Continuous recurrent neural networks with adversarial training
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
- Audio Spectrogram Representations for Processing with Convolutional Neural Networks
- MelNet: A Generative Model for Audio in the Frequency Domain
- Jukebox: A Generative Model for Music
- High Fidelity Speech Synthesis with Adversarial Networks
- Counterpoint by Convolution
- DDSP: Differentiable Digital Signal Processing
- POP909: A Pop-song Dataset for Music Arrangement Generation
- Latent Normalizing Flows for Discrete Sequences
- MMM : Exploring Conditional Multi-Track Music Generation with the Transformer
- Song From PI: A Musically Plausible Network for Pop Music Generation
- Singing voice synthesis based on convolutional neural networks
- Deep Long Audio Inpainting
- Interactive Music Generation with Positional Constraints using Anticipation-RNNs
- PIANOTREE VAE: Structured Representation Learning for Polyphonic Music
- Neural Translation of Musical Style
- XiaoiceSing: A High-Quality and Integrated Singing Voice Synthesis System
- DurIAN-SC: Duration Informed Attention Network based Singing Voice Conversion System
- Vector Quantized Contrastive Predictive Coding for Template-based Music Generation
- Explicitly Conditioned Melody Generation: A Case Study with Interdependent RNNs
- Time Domain Neural Audio Style Transfer
- Learning Singing From Speech
- NONOTO: A Model-agnostic Web Interface for Interactive Music Composition by Inpainting
- Adversarially Trained Multi-Singer Sequence-To-Sequence Singing Synthesizer
- Learning Style-Aware Symbolic Music Representations by Adversarial Autoencoders
- Inspecting and Interacting with Meaningful Music Representations using VAE
- Music Generation with Deep Learning
- Peking Opera Synthesis via Duration Informed Attention Network
- Generative Modelling for Controllable Audio Synthesis of Expressive Piano Performance
- DeepDrummer : Generating Drum Loops using Deep Learning and a Human in the Loop