105 citations · 165 across the 12 of their papers we have counts for
12 papers
Multi-view and Multi-modal Event Detection Utilizing Transformer-based Multi-sensor fusion
Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito +1
We tackle a challenging task: multi-view and multi-modal event detection that detects events in a wide-range real environment by utilizing data from distributed cameras and microph…
Echo-aware Adaptation of Sound Event Localization and Detection in Unknown Environments
Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito
Our goal is to develop a sound event localization and detection (SELD) system that works robustly in unknown environments. A SELD system trained on known environment data is degrad…
Wearable SELD dataset: Dataset for sound event localization and detection using wearable devices around head
Kento Nagatomo, Masahiro Yasuda, Kohei Yatabe +2
Sound event localization and detection (SELD) is a combined task of identifying the sound event and its direction. Deep neural networks (DNNs) are utilized to associate them with t…
APPLADE: Adjustable Plug-and-play Audio Declipper Combining DNN with Sparse Optimization
Tomoro Tanaka, Kohei Yatabe, Masahiro Yasuda +1
In this paper, we propose an audio declipping method that takes advantages of both sparse optimization and deep learning. Since sparsity-based audio declipping methods have been de…
ToyADMOS2: Another dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions
Noboru Harada, Daisuke Niizumi, Daiki Takeuchi +3
This paper proposes a new large-scale dataset called "ToyADMOS2" for anomaly detection in machine operating sounds (ADMOS). As did for our previous ToyADMOS dataset, we collected a…
Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval
Yuma Koizumi, Yasunori Ohishi, Daisuke Niizumi +2
The goal of audio captioning is to translate input audio into its description using natural language. One of the problems in audio captioning is the lack of training data due to th…