19 citations · 54 across the 9 of their papers we have counts for
15 papers
Ensemble of ACCDOA- and EINV2-based Systems with D3Nets and Impulse Response Simulation for Sound Event Localization and Detection
Kazuki Shimada, Naoya Takahashi, Yuichiro Koyama +4
This report describes our systems submitted to the DCASE2021 challenge task 3: sound event localization and detection (SELD) with directional interference. Our previous system base…
Training Speech Enhancement Systems with Noisy Speech Datasets
Koichi Saito, Stefan Uhlich, Giorgio Fabbro +1
Recently, deep neural network (DNN)-based speech enhancement (SE) systems have been used with great success. During training, such systems require clean speech data - ideally, in l…
Hierarchical disentangled representation learning for singing voice conversion
Naoya Takahashi, Mayank Kumar Singh, Yuki Mitsufuji
Conventional singing voice conversion (SVC) methods often suffer from operating in high-resolution audio owing to a high dimensionality of data. In this paper, we propose a hierarc…
Densely connected multidilated convolutional networks for dense prediction tasks
Naoya Takahashi, Yuki Mitsufuji
Tasks that involve high-resolution dense prediction require a modeling of both local and global patterns in a large input field. Although the local and global structures often depe…
ACCDOA: Activity-Coupled Cartesian Direction of Arrival Representation for Sound Event Localization and Detection
Kazuki Shimada, Yuichiro Koyama, Naoya Takahashi +2
Neural-network (NN)-based methods show high performance in sound event localization and detection (SELD). Conventional NN-based methods use two branches for a sound event detection…
Adversarial attacks on audio source separation
Naoya Takahashi, Shota Inoue, Yuki Mitsufuji
Despite the excellent performance of neural-network-based audio source separation methods and their wide range of applications, their robustness against intentional attacks has bee…