A Tutorial on Deep Learning for Music Information Retrieval
arXiv:1709.04396
Abstract
Following their success in Computer Vision and other areas, deep learning techniques have recently become widely adopted in Music Information Retrieval (MIR) research. However, the majority of works aim to adopt and assess methods that have been shown to be effective in other domains, while there is still a great need for more original research focusing on music primarily and utilising musical knowledge and insight. The goal of this paper is to boost the interest of beginners by providing a comprehensive tutorial and reducing the barriers to entry into deep learning for MIR. We lay out the basic principles and review prominent works in this hard to navigate the field. We then outline the network structures that have been successful in MIR problems and facilitate the selection of building blocks for the problems at hand. Finally, guidelines for new tasks and some advanced topics in deep learning are discussed to stimulate new research in this fascinating field.
References in corpus (23)
- Deep Learning in Neural Networks: An Overview
- Sequence to Sequence Learning with Neural Networks
- Conditional Generative Adversarial Nets
- WaveNet: A Generative Model for Raw Audio
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- BEGAN: Boundary Equilibrium Generative Adversarial Networks
- Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
- MidiNet: A Convolutional Generative Adversarial Network for Symbolic-domain Music Generation
- SoundNet: Learning Sound Representations from Unlabeled Video
- Blocks and Fuel: Frameworks for deep learning
- Sample-level Deep Convolutional Neural Networks for Music Auto-tagging Using Raw Waveforms
- A Fully Convolutional Deep Auditory Model for Musical Chord Recognition
- On the Potential of Simple Framewise Approaches to Piano Transcription
- Feature Learning for Chord Recognition: The Deep Chroma Extractor
- Kapre: On-GPU Audio Preprocessing Layers for a Quick Implementation of Deep Neural Network Models with Keras
- Stacked Convolutional and Recurrent Neural Networks for Music Emotion Recognition
- Explaining Deep Convolutional Neural Networks on Music Classification
- Deep Karaoke: Extracting Vocals from Musical Mixtures Using a Convolutional Deep Neural Network
- Timbre Analysis of Music Audio Signals with Convolutional Neural Networks
- Deep Cross-Modal Audio-Visual Generation
- Revisiting the problem of audio-based hit song prediction using convolutional neural networks
- The Effects of Noisy Labels on Deep Convolutional Neural Networks for Music Tagging
Cited by in corpus (9)
- Audio Spoofing Verification using Deep Convolutional Neural Networks by Transfer Learning
- Pop Music Highlighter: Marking the Emotion Keypoints
- What all do audio transformer models hear? Probing Acoustic Representations for Language Delivery and its Structure
- DLR : Toward a deep learned rhythmic representation for music content analysis
- Exploiting Synchronized Lyrics And Vocal Features For Music Emotion Detection
- The Effects of Noisy Labels on Deep Convolutional Neural Networks for Music Tagging
- JTAV: Jointly Learning Social Media Content Representation by Fusing Textual, Acoustic, and Visual Features
- Ultra-light deep MIR by trimming lottery tickets
- MusicTM-Dataset for Joint Representation Learning among Sheet Music, Lyrics, and Musical Audio