9 papers
Description and Discussion on DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
Binh Thien Nguyen, Masahiro Yasuda, Noboru Harada +8
This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 4, Spatial Semantic Segmentation of Sound Scenes (S5).…
DiBA: Diagonal and Binary Matrix Approximation for Neural Network Weight Compression
Nobutaka Ono
In this paper, we propose DiBA (Diagonal and Binary Matrix Approximation), a compact matrix factorization for neural network weight compression. Many components of modern networks,…
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda +3
Since the introduction of Masked Autoencoders, various improvements to masking techniques have been explored. In this paper, we rethink masking strategies for audio representation…
What Do Neurons Listen To? A Neuron-level Dissection of a General-purpose Audio Model
Takao Kawamura, Daisuke Niizumi, Nobutaka Ono
In this paper, we analyze the internal representations of a general-purpose audio self-supervised learning (SSL) model from a neuron-level perspective. Despite their strong empiric…
Fast Swap-Based Element Selection for Multiplication-Free Dimension Reduction
Nobutaka Ono
In this paper, we propose a fast algorithm for element selection, a multiplication-free form of dimension reduction that produces a dimension-reduced vector by simply selecting a s…
On the Invariance of Cross-Correlation Peak Positions Under Monotonic Signal Transformations, with Application to Fast Time Difference Estimation
Natsuki Ueno, Ryotaro Sato, Nobutaka Ono
We present a theorem concerning the invariance of cross-correlation peak positions. This theoretical result provides the foundation for a new method for time difference estimation…