Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano Compositions
arXiv:2002.00212
Abstract
A great number of deep learning based models have been recently proposed for automatic music composition. Among these models, the Transformer stands out as a prominent approach for generating expressive classical piano performance with a coherent structure of up to one minute. The model is powerful in that it learns abstractions of data on its own, without much human-imposed domain knowledge or constraints. In contrast with this general approach, this paper shows that Transformers can do even better for music modeling, when we improve the way a musical score is converted into the data fed to a Transformer model. In particular, we seek to impose a metrical structure in the input data, so that Transformers can be more easily aware of the beat-bar-phrase hierarchical structure in music. The new data representation maintains the flexibility of local tempo changes, and provides hurdles to control the rhythmic and harmonic structure of music. With this approach, we build a Pop Music Transformer that composes Pop piano music with better rhythmic structure than existing Transformer models.
Accepted at ACM Multimedia 2020
References in corpus (3)
Cited by in corpus (12)
- Video Background Music Generation with Controllable Music Transformer
- Unconditional Audio Generation with Generative Adversarial Networks and Cycle Regularization
- MusPy: A Toolkit for Symbolic Music Generation
- A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music Modeling
- Controllable deep melody generation via hierarchical music structure representation
- Dual Learning Music Composition and Dance Choreography
- PopMAG: Pop Music Accompaniment Generation
- The Beauty of Repetition in Machine Composition Scenarios
- Melody Structure Transfer Network: Generating Music with Separable Self-Attention
- Style-based Composer Identification and Attribution of Symbolic Music Scores: a Systematic Survey
- Exploring Inherent Properties of the Monophonic Melody of Songs
- MuSLCAT: Multi-Scale Multi-Level Convolutional Attention Transformer for Discriminative Music Modeling on Raw Waveforms