Sound Event Detection and Separation: a Benchmark on Desed Synthetic Soundscapes
arXiv:2011.00801
Abstract
We propose a benchmark of state-of-the-art sound event detection systems (SED). We designed synthetic evaluation sets to focus on specific sound event detection challenges. We analyze the performance of the submissions to DCASE 2021 task 4 depending on time related modifications (time position of an event and length of clips) and we study the impact of non-target sound events and reverberation. We show that the localization in time of sound events is still a problem for SED systems. We also show that reverberation and non-target sound events are severely degrading the performance of the SED systems. In the latter case, sound separation seems like a promising solution.
Cited by in corpus (4)
- Computational bioacoustics with deep learning: a review and roadmap
- You Only Hear Once: A YOLO-like Algorithm for Audio Segmentation and Sound Event Detection
- Self-training with noisy student model and semi-supervised loss function for dcase 2021 challenge task 4
- A benchmark of state-of-the-art sound event detection systems evaluated on synthetic soundscapes