105 citations · 280 across the 31 of their papers we have counts for
12 papers · 1 filter
Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval
Yuma Koizumi, Yasunori Ohishi, Daisuke Niizumi +2
The goal of audio captioning is to translate input audio into its description using natural language. One of the problems in audio captioning is the lack of training data due to th…
Effects of Word-frequency based Pre- and Post- Processings for Audio Captioning
Daiki Takeuchi, Yuma Koizumi, Yasunori Ohishi +2
The system we used for Task 6 (Automated Audio Captioning)of the Detection and Classification of Acoustic Scenes and Events(DCASE) 2020 Challenge combines three elements, namely, d…
A Transformer-based Audio Captioning Model with Keyword Estimation
Yuma Koizumi, Ryo Masumura, Kyosuke Nishida +2
One of the problems with automated audio captioning (AAC) is the indeterminacy in word selection corresponding to the audio event/scene. Since one acoustic event/scene can be descr…
The NTT DCASE2020 Challenge Task 6 system: Automated Audio Captioning with Keywords and Sentence Length Estimation
Yuma Koizumi, Daiki Takeuchi, Yasunori Ohishi +2
This technical report describes the system participating to the Detection and Classification of Acoustic Scenes and Events (DCASE) 2020 Challenge, Task 6: automated audio captionin…
Listen to What You Want: Neural Network-based Universal Sound Selector
Tsubasa Ochiai, Marc Delcroix, Yuma Koizumi +3
Being able to control the acoustic events (AEs) to which we want to listen would allow the development of more controllable hearable devices. This paper addresses the AE sound sele…
Description and Discussion on DCASE2020 Challenge Task2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
Yuma Koizumi, Yohei Kawaguchi, Keisuke Imoto +8
In this paper, we present the task description and discuss the results of the DCASE 2020 Challenge Task 2: Unsupervised Detection of Anomalous Sounds for Machine Condition Monitori…