25 citations · 96 across the 10 of their papers we have counts for
10 papers · 1 filter
Composing General Audio Representation by Fusing Multilayer Features of a Pre-trained Model
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2
Many application studies rely on audio DNN models pre-trained on a large-scale dataset as essential feature extractors, and they extract features from the last layers. In this stud…
ToyADMOS2: Another dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions
Noboru Harada, Daisuke Niizumi, Daiki Takeuchi +3
This paper proposes a new large-scale dataset called "ToyADMOS2" for anomaly detection in machine operating sounds (ADMOS). As did for our previous ToyADMOS dataset, we collected a…
BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2
Inspired by the recent progress in self-supervised learning for computer vision that generates supervision using data augmentations, we explore a new general-purpose audio represen…
Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval
Yuma Koizumi, Yasunori Ohishi, Daisuke Niizumi +2
The goal of audio captioning is to translate input audio into its description using natural language. One of the problems in audio captioning is the lack of training data due to th…
Effects of Word-frequency based Pre- and Post- Processings for Audio Captioning
Daiki Takeuchi, Yuma Koizumi, Yasunori Ohishi +2
The system we used for Task 6 (Automated Audio Captioning)of the Detection and Classification of Acoustic Scenes and Events(DCASE) 2020 Challenge combines three elements, namely, d…
The NTT DCASE2020 Challenge Task 6 system: Automated Audio Captioning with Keywords and Sentence Length Estimation
Yuma Koizumi, Daiki Takeuchi, Yasunori Ohishi +2
This technical report describes the system participating to the Detection and Classification of Acoustic Scenes and Events (DCASE) 2020 Challenge, Task 6: automated audio captionin…