43 citations · 96 across the 12 of their papers we have counts for
8 papers · 1 filter
Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification
Kazuki Shimada, Archontis Politis, Iran R. Roman +10
This paper presents the objective, dataset, baseline, and metrics of Task 3 of the DCASE2025 Challenge on sound event localization and detection (SELD). In previous editions, the c…
Music Foundation Model as Generic Booster for Music Downstream Tasks
WeiHsiang Liao, Yuhta Takida, Yukara Ikemiya +13
We demonstrate the efficacy of using intermediate representations from a single foundation model to enhance various music downstream tasks. We introduce SoniDo, a music foundation…
Zero- and Few-shot Sound Event Localization and Detection
Kazuki Shimada, Kengo Uchida, Yuichiro Koyama +4
Sound event localization and detection (SELD) systems estimate direction-of-arrival (DOA) and temporal activation for sets of target classes. Neural network (NN)-based SELD systems…
STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events
Kazuki Shimada, Archontis Politis, Parthasaarathy Sudarsanam +9
While direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perce…
Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
Hao Shi, Kazuki Shimada, Masato Hirano +6
Diffusion-based generative speech enhancement (SE) has recently received attention, but reverse diffusion remains time-consuming. One solution is to initialize the reverse diffusio…
Spatial Data Augmentation with Simulated Room Impulse Responses for Sound Event Localization and Detection
Yuichiro Koyama, Kazuhide Shigemi, Masafumi Takahashi +5
Recording and annotating real sound events for a sound event localization and detection (SELD) task is time consuming, and data augmentation techniques are often favored when the a…