Deep convolutional networks on the pitch spiral for musical instrument recognition
arXiv:1605.06644
Abstract
Musical performance combines a wide range of pitches, nuances, and expressive techniques. Audio-based classification of musical instruments thus requires to build signal representations that are invariant to such transformations. This article investigates the construction of learned convolutional architectures for instrument recognition, given a limited amount of annotated training data. In this context, we benchmark three different weight sharing strategies for deep convolutional networks in the time-frequency domain: temporal kernels; time-frequency kernels; and a linear combination of time-frequency kernels which are one octave apart, akin to a Shepard pitch spiral. We provide an acoustical interpretation of these strategies within the source-filter framework of quasi-harmonic sounds with a fixed spectral envelope, which are archetypal of musical notes. The best classification accuracy is obtained by hybridizing all three convolutional layers into a single deep learning architecture.
7 pages, 3 figures. Accepted at the International Society for Music Information Retrieval Conference (ISMIR) conference in New York City, NY, USA, August 2016
References in corpus (1)
Cited by in corpus (9)
- A Tutorial on Deep Learning for Music Information Retrieval
- Wavelet Moments for Cosmological Parameter Estimation
- Lyrics-Based Music Genre Classification Using a Hierarchical Attention Network
- Extended playing techniques: The next milestone in musical instrument recognition
- Learning a Lie Algebra from Unlabeled Data Pairs
- One or Two Components? The Scattering Transform Answers
- Visual Attention for Musical Instrument Recognition
- Helicality: An Isomap-based Measure of Octave Equivalence in Audio Data
- Machine listening intelligence