End-to-end learning for music audio tagging at scale
arXiv:1711.02520
Abstract
The lack of data tends to limit the outcomes of deep learning research, particularly when dealing with end-to-end learning stacks processing raw data such as waveforms. In this study, 1.2M tracks annotated with musical labels are available to train our end-to-end models. This large amount of data allows us to unrestrictedly explore two different design paradigms for music auto-tagging: assumption-free models - using waveforms as input with very small convolutional filters; and models that rely on domain knowledge - log-mel spectrograms with a convolutional neural network designed to learn timbral and temporal features. Our work focuses on studying how these two types of deep architectures perform when datasets of variable size are available for training: the MagnaTagATune (25k songs), the Million Song Dataset (240k songs), and a private dataset of 1.2M songs. Our experiments suggest that music domain assumptions are relevant when not enough training data are available, thus showing how waveform-based models outperform spectrogram-based ones in large-scale data scenarios.
Presented at the Workshop on Machine Learning for Audio Signal Processing (ML4Audio) at NIPS 2017, and in proceedings of the 19th International Society for Music Information Retrieval Conference (ISMIR2018). Code: https://github.com/jordipons/music-audio-tagging-at-scale-models. Demo: http://www.jordipons.me/apps/music-audio-tagging-at-scale-demo/
References in corpus (3)
Cited by in corpus (14)
- GACELA -- A generative adversarial context encoder for long audio inpainting
- Explaining Deep Classification of Time-Series Data with Learned Prototypes
- MusiCoder: A Universal Music-Acoustic Encoder Based on Transformers
- Learning to rank music tracks using triplet loss
- Adversarial Generation of Time-Frequency Features with application in audio synthesis
- Learning Music Sequence Representation from Text Supervision
- Artificial Musical Intelligence: A Survey
- nnAudio: An on-the-fly GPU Audio to Spectrogram Conversion Toolbox Using 1D Convolution Neural Networks
- Improving Machine Hearing on Limited Data Sets
- A Feature Learning Siamese Model for Intelligent Control of the Dynamic Range Compressor
- Modeling of nonlinear audio effects with end-to-end deep neural networks
- A Hybrid Approach to Audio-to-Score Alignment
- Towards Learning Universal Audio Representations
- Multi-scale Embedded CNN for Music Tagging (MsE-CNN)