5 papers
Input-Adaptive Spectral Feature Compression by Sequence Modeling for Source Separation
Kohei Saijo, Yoshiaki Bando
Time-frequency domain dual-path models have demonstrated strong performance and are widely used in source separation. Because their computational cost grows with the number of freq…
Is MixIT Really Unsuitable for Correlated Sources? Exploring MixIT for Unsupervised Pre-training in Music Source Separation
Kohei Saijo, Yoshiaki Bando
In music source separation (MSS), obtaining isolated sources or stems is highly costly, making pre-training on unlabeled data a promising approach. Although source-agnostic unsuper…
Formula-Supervised Sound Event Detection: Pre-Training Without Real Data
Yuto Shibata, Keitaro Tanaka, Yoshiaki Bando +3
In this paper, we propose a novel formula-driven supervised learning (FDSL) framework for pre-training an environmental sound analysis model by leveraging acoustic signals parametr…
DOA-Aware Audio-Visual Self-Supervised Learning for Sound Event Localization and Detection
Yoto Fujita, Yoshiaki Bando, Keisuke Imoto +2
This paper describes sound event localization and detection (SELD) for spatial audio recordings captured by firstorder ambisonics (FOA) microphones. In this task, one may train a d…
LEAD Dataset: How Can Labels for Sound Event Detection Vary Depending on Annotators?
Naoki Koga, Yoshiaki Bando, Keisuke Imoto
In this paper, we introduce a LargE-scale Annotator's labels for sound event Detection (LEAD) dataset, which is the dataset used to gain a better understanding of the variation in…