A Sequence Matching Network for Polyphonic Sound Event Localization and Detection
arXiv:2002.05865 · doi:10.1109/ICASSP40776.2020.9053045
Abstract
Polyphonic sound event detection and direction-of-arrival estimation require different input features from audio signals. While sound event detection mainly relies on time-frequency patterns, direction-of-arrival estimation relies on magnitude or phase differences between microphones. Previous approaches use the same input features for sound event detection and direction-of-arrival estimation, and train the two tasks jointly or in a two-stage transfer-learning manner. We propose a two-step approach that decouples the learning of the sound event detection and directional-of-arrival estimation systems. In the first step, we detect the sound events and estimate the directions-of-arrival separately to optimize the performance of each system. In the second step, we train a deep neural network to match the two output sequences of the event detector and the direction-of-arrival estimator. This modular and hierarchical approach allows the flexibility in the system design, and increase the performance of the whole sound event localization and detection system. The experimental results using the DCASE 2019 sound event localization and detection dataset show an improved performance compared to the previous state-of-the-art solutions.
to be published in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
References in corpus (1)
Cited by in corpus (11)
- A Survey of Sound Source Localization with Deep Learning Methods
- Overview and Evaluation of Sound Event Localization and Detection in DCASE 2019
- SALSA: Spatial Cue-Augmented Log-Spectrogram Features for Polyphonic Sound Event Localization and Detection
- SALSA-Lite: A Fast and Effective Feature for Polyphonic Sound Event Localization and Detection with Microphone Arrays
- Event-Independent Network for Polyphonic Sound Event Localization and Detection
- A Dataset of Dynamic Reverberant Sound Scenes with Directional Interferers for Sound Event Localization and Detection
- ACCDOA: Activity-Coupled Cartesian Direction of Arrival Representation for Sound Event Localization and Detection
- An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection
- Leveraging Visual Supervision for Array-based Active Speaker Detection and Localization
- A General Network Architecture for Sound Event Localization and Detection Using Transfer Learning and Recurrent Neural Network
- What Makes Sound Event Localization and Detection Difficult? Insights from Error Analysis