DCASE 2018 Challenge Surrey Cross-Task convolutional neural network baseline
arXiv:1808.00773
Abstract
The Detection and Classification of Acoustic Scenes and Events (DCASE) consists of five audio classification and sound event detection tasks: 1) Acoustic scene classification, 2) General-purpose audio tagging of Freesound, 3) Bird audio detection, 4) Weakly-labeled semi-supervised sound event detection and 5) Multi-channel audio classification. In this paper, we create a cross-task baseline system for all five tasks based on a convlutional neural network (CNN): a "CNN Baseline" system. We implemented CNNs with 4 layers and 8 layers originating from AlexNet and VGG from computer vision. We investigated how the performance varies from task to task with the same configuration of neural networks. Experiments show that deeper CNN with 8 layers performs better than CNN with 4 layers on all tasks except Task 1. Using CNN with 8 layers, we achieve an accuracy of 0.680 on Task 1, an accuracy of 0.895 and a mean average precision (MAP) of 0.928 on Task 2, an accuracy of 0.751 and an area under the curve (AUC) of 0.854 on Task 3, a sound event detection F1 score of 20.8% on Task 4, and an F1 score of 87.75% on Task 5. We released the Python source code of the baseline systems under the MIT license for further research.
Accepted by DCASE 2018 Workshop. 4 pages. Source code available
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Convolutional Recurrent Neural Networks for Polyphonic Sound Event Detection
- Automatic acoustic detection of birds through deep learning: the first Bird Audio Detection challenge
- SoundNet: Learning Sound Representations from Unlabeled Video
- A multi-device dataset for urban acoustic scene classification
- General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline
- Large-scale weakly supervised audio classification using gated convolutional neural network
- A Comparison of deep learning methods for environmental sound
Cited by in corpus (7)
- A Squeeze-and-Excitation and Transformer based Cross-task System for Environmental Sound Recognition
- Sound Context Classification Basing on Join Learning Model and Multi-Spectrogram Features
- CNN depth analysis with different channel inputs for Acoustic Scene Classification
- HODGEPODGE: Sound event detection based on ensemble of semi-supervised learning methods
- DD-CNN: Depthwise Disout Convolutional Neural Network for Low-complexity Acoustic Scene Classification
- Cross-modal Spectrum Transformation Network For Acoustic Scene classification
- Hodge and Podge: Hybrid Supervised Sound Event Detection with Multi-Hot MixMatch and Composition Consistence Training