859 citations · 1.5k across the 31 of their papers we have counts for
23 papers · 1 filter
Language-based Audio Retrieval Task in DCASE 2022 Challenge
Huang Xie, Samuel Lipping, Tuomas Virtanen
Language-based audio retrieval is a task, where natural language textual captions are used as queries to retrieve audio signals from a dataset. It has been first introduced into DC…
Differentiable Tracking-Based Training of Deep Learning Sound Source Localizers
Sharath Adavanne, Archontis Politis, Tuomas Virtanen
Data-based and learning-based sound source localization (SSL) has shown promising results in challenging conditions, and is commonly set as a classification or a regression problem…
Sound Event Detection: A Tutorial
Annamaria Mesaros, Toni Heittola, Tuomas Virtanen +1
The goal of automatic sound event detection (SED) methods is to recognize what is happening in an audio signal and when it is happening. In practice, the goal is to recognize at wh…
A Dataset of Dynamic Reverberant Sound Scenes with Directional Interferers for Sound Event Localization and Detection
Archontis Politis, Sharath Adavanne, Daniel Krause +3
This report presents the dataset and baseline of Task 3 of the DCASE2021 Challenge on Sound Event Localization and Detection (SELD). The dataset is based on emulation of real recor…
Mobile Microphone Array Speech Detection and Localization in Diverse Everyday Environments
Pasi Pertilä, Emre Cakir, Aapo Hakala +4
Joint sound event localization and detection (SELD) is an integral part of developing context awareness into communication interfaces of mobile robots, smartphones, and home assist…
Deep neural network Based Low-latency Speech Separation with Asymmetric analysis-Synthesis Window Pair
Shanshan Wang, Gaurav Naithani, Archontis Politis +1
Time-frequency masking or spectrum prediction computed via short symmetric windows are commonly used in low-latency deep neural network (DNN) based source separation. In this paper…