11 citations · 24 across the 20 of their papers we have counts for
8 papers · 1 filter
Description and Discussion on DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
Binh Thien Nguyen, Masahiro Yasuda, Noboru Harada +8
This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 4, Spatial Semantic Segmentation of Sound Scenes (S5).…
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda +3
Since the introduction of Masked Autoencoders, various improvements to masking techniques have been explored. In this paper, we rethink masking strategies for audio representation…
What Do Neurons Listen To? A Neuron-level Dissection of a General-purpose Audio Model
Takao Kawamura, Daisuke Niizumi, Nobutaka Ono
In this paper, we analyze the internal representations of a general-purpose audio self-supervised learning (SSL) model from a neuron-level perspective. Despite their strong empiric…
Incremental Averaging Method to Improve Graph-Based Time-Difference-of-Arrival Estimation
Klaus Brümann, Kouei Yamaoka, Nobutaka Ono +1
Estimating the position of a speech source based on time-differences-of-arrival (TDOAs) is often adversely affected by background noise and reverberation. A popular method to estim…
Mel-Spectrogram Inversion via Alternating Direction Method of Multipliers
Yoshiki Masuyama, Natsuki Ueno, Nobutaka Ono
Signal reconstruction from its mel-spectrogram is known as mel-spectrogram inversion and has many applications, including speech and foley sound synthesis. In this paper, we propos…
Joint Dereverberation and Separation with Iterative Source Steering
Taishi Nakashima, Robin Scheibler, Masahito Togami +1
We propose a new algorithm for joint dereverberation and blind source separation (DR-BSS). Our work builds upon the IRLMA-T framework that applies a unified filter combining dereve…