Deep Convolutional Neural Networks and Data Augmentation for Acoustic Event Detection
arXiv:1604.07160
Abstract
We propose a novel method for Acoustic Event Detection (AED). In contrast to speech, sounds coming from acoustic events may be produced by a wide variety of sources. Furthermore, distinguishing them often requires analyzing an extended time period due to the lack of a clear sub-word unit. In order to incorporate the long-time frequency structure for AED, we introduce a convolutional neural network (CNN) with a large input field. In contrast to previous works, this enables to train audio event detection end-to-end. Our architecture is inspired by the success of VGGNet and uses small, 3x3 convolutions, but more depth than previous methods in AED. In order to prevent over-fitting and to take full advantage of the modeling capabilities of our network, we further propose a novel data augmentation method to introduce data variation. Experimental results show that our CNN significantly outperforms state of the art methods including Bag of Audio Words (BoAW) and classical CNNs, achieving a 16% absolute improvement.
Presented in INTERSPEECH 2016
References in corpus (1)
Cited by in corpus (15)
- CUHK & ETHZ & SIAT Submission to ActivityNet Challenge 2016
- Environmental Sound Classification on the Edge: A Pipeline for Deep Acoustic Networks on Extremely Resource-Constrained Devices
- Pruning vs XNOR-Net: A Comprehensive Study of Deep Learning for Audio Classification on Edge-devices
- A Four-Stage Data Augmentation Approach to ResNet-Conformer Based Acoustic Modeling for Sound Event Localization and Detection
- Ensemble of ACCDOA- and EINV2-based Systems with D3Nets and Impulse Response Simulation for Sound Event Localization and Detection
- ACCDOA: Activity-Coupled Cartesian Direction of Arrival Representation for Sound Event Localization and Detection
- Sound Event Localization and Detection Using Activity-Coupled Cartesian DOA Vector and RD3net
- Attention based Convolutional Recurrent Neural Network for Environmental Sound Classification
- SwishNet: A Fast Convolutional Neural Network for Speech, Music and Noise Classification and Segmentation
- Multi-Temporal Resolution Convolutional Neural Networks for Acoustic Scene Classification
- Unsupervised Learning of Semantic Audio Representations
- AI for Earth: Rainforest Conservation by Acoustic Surveillance
- Data Augmentation via Dependency Tree Morphing for Low-Resource Languages
- Deep Convolutional Neural Network with Mixup for Environmental Sound Classification
- An End-to-End Audio Classification System based on Raw Waveforms and Mix-Training Strategy