3 citations · 3 across the 3 of their papers we have counts for
8 papers
Quantifying the Uncertainty of Blindly Estimated Room Embeddings Using a Dispersion-Calibrated Score
Yang Xiang, Philipp Götz, Emanuël A. P. Habets +3
Room embeddings derived from reverberant speech are often unreliable: speech content and recording degradation can alter the representation even when speaker, room, and source-rece…
Grammar-Guided Hierarchical Parsing for Long-form Audio Activity Recognition
Peng Zhang, Qingyu Luo, Philip J. B. Jackson +1
Long-form audio exhibits an inherent hierarchy: fine-grained events form sub-activities, which in turn constitute higher-level activities. Prior work often models these levels sepa…
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
Davide Berghi, Philip J. B. Jackson
Sound event localization and detection (SELD) involves predicting active sound event classes over time while estimating their positions. The localization subtask in SELD is usually…
ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video
Davide Berghi, Philip J. B. Jackson
Sound event localization and detection with distance estimation (3D SELD) in video involves identifying active sound events at each time frame while estimating their spatial coordi…
Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
Davide Berghi, Philip J. B. Jackson
In this study, we address the multimodal task of stereo sound event localization and detection with source distance estimation (3D SELD) in regular video content. 3D SELD is a comp…
Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos
Davide Berghi, Philip J. B. Jackson
This report presents our systems submitted to the audio-only and audio-visual tracks of the DCASE2025 Task 3 Challenge: Stereo Sound Event Localization and Detection (SELD) in Regu…