Recurrent Neural Networks for Polyphonic Sound Event Detection in Real Life Recordings
arXiv:1604.00861 · doi:10.1109/ICASSP.2016.7472917
Abstract
In this paper we present an approach to polyphonic sound event detection in real life recordings based on bi-directional long short term memory (BLSTM) recurrent neural networks (RNNs). A single multilabel BLSTM RNN is trained to map acoustic features of a mixture signal consisting of sounds from multiple classes, to binary activity indicators of each event class. Our method is tested on a large database of real-life recordings, with 61 classes (e.g. music, car, speech) from 10 different everyday contexts. The proposed method outperforms previous approaches by a large margin, and the results are further improved using data augmentation techniques. Overall, our system reports an average F1-score of 65.5% on 1 second blocks and 64.7% on single frames, a relative improvement over previous state-of-the-art approach of 6.8% and 15.1% respectively.
To appean in Proceedings of IEEE ICASSP 2016
Cited by in corpus (41)
- Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification
- Deep Learning for Audio Signal Processing
- Convolutional Recurrent Neural Networks for Polyphonic Sound Event Detection
- Sound Event Localization and Detection of Overlapping Sources Using Convolutional Recurrent Neural Networks
- Sound Event Detection: A Tutorial
- Polyphonic Sound Event Detection and Localization using a Two-Stage Strategy
- Benchmarking of eight recurrent neural network variants for breath phase and adventitious sound detection on a self-developed open-access lung sound database-HF_Lung_V1
- Sound Event Detection in Multichannel Audio Using Spatial and Harmonic Features
- Sound Event Detection and Time-Frequency Segmentation from Weakly Labelled Data
- A report on sound event detection with different binaural features
- A Sequence Matching Network for Polyphonic Sound Event Localization and Detection
- DNN and CNN with Weighted and Multi-task Loss Functions for Audio Event Detection
- AI-based soundscape analysis: Jointly identifying sound sources and predicting annoyance
- What Makes Audio Event Detection Harder than Classification?
- A Four-Stage Data Augmentation Approach to ResNet-Conformer Based Acoustic Modeling for Sound Event Localization and Detection
- Sound Event Detection with Sequentially Labelled Data Based on Connectionist Temporal Classification and Unsupervised Clustering
- On the Benefits of Early Fusion in Multimodal Representation Learning
- Efficient Convolutional Neural Network For Audio Event Detection
- A Multi-grained based Attention Network for Semi-supervised Sound Event Detection
- Enabling Early Audio Event Detection with Neural Networks
- Unified Multisensory Perception: Weakly-Supervised Audio-Visual Video Parsing
- Acoustic Event Detection with Classifier Chains
- VGGSound: A Large-scale Audio-Visual Dataset
- Adaptive pooling operators for weakly labeled sound event detection
- Sound event detection via dilated convolutional recurrent neural networks
- EnvGAN: Adversarial Synthesis of Environmental Sounds for Data Augmentation
- GISE-51: A scalable isolated sound events dataset
- Weakly Labeled Sound Event Detection Using Tri-training and Adversarial Learning
- Specialized Decision Surface and Disentangled Feature for Weakly-Supervised Polyphonic Sound Event Detection
- An Ensemble Framework of Voice-Based Emotion Recognition System for Films and TV Programs
- A Comparison of Pooling Methods on LSTM Models for Rare Acoustic Event Classification
- A Capsule based Approach for Polyphonic Sound Event Detection
- Enhancing Sound Texture in CNN-Based Acoustic Scene Classification
- A two-step system for sound event localization and detection
- Active Learning for Sound Event Detection
- Guided learning for weakly-labeled semi-supervised sound event detection
- Transferring Voice Knowledge for Acoustic Event Detection: An Empirical Study
- Environment Transfer for Distributed Systems
- Adaptive Multi-scale Detection of Acoustic Events
- Intra-Utterance Similarity Preserving Knowledge Distillation for Audio Tagging
- Power pooling: An adaptive pooling function for weakly labelled sound event detection