most citedThe NTT DCASE2020 Challenge Task 6 system: Automated Audio Captioning with Keywords and Sentence Length Estimation

21 citations · 49 across the 6 of their papers we have counts for

collaborators

6 papers

eess.AS202011 cited

Effects of Word-frequency based Pre- and Post- Processings for Audio Captioning

Daiki Takeuchi, Yuma Koizumi, Yasunori Ohishi +2

The system we used for Task 6 (Automated Audio Captioning)of the Detection and Classification of Acoustic Scenes and Events(DCASE) 2020 Challenge combines three elements, namely, d…

eess.AS202021 cited

The NTT DCASE2020 Challenge Task 6 system: Automated Audio Captioning with Keywords and Sentence Length Estimation

Yuma Koizumi, Daiki Takeuchi, Yasunori Ohishi +2

This technical report describes the system participating to the Detection and Classification of Acoustic Scenes and Events (DCASE) 2020 Challenge, Task 6: automated audio captionin…

eess.AS20208 cited

Speech Enhancement using Self-Adaptation and Multi-Head Self-Attention

Yuma Koizumi, Kohei Yatabe, Marc Delcroix +2

This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly fro…

eess.AS20205 cited

Real-time speech enhancement using equilibriated RNN

Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi +2

We propose a speech enhancement method using a causal deep neural network~(DNN) for real-time applications. DNN has been widely used for estimating a time-frequency~(T-F) mask whic…

eess.AS20191 cited

Invertible DNN-based nonlinear time-frequency transform for speech enhancement

Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi +2

We propose an end-to-end speech enhancement method with trainable time-frequency~(T-F) transform based on invertible deep neural network~(DNN). The resent development of speech enh…

eess.AS20193 cited

Data-driven design of perfect reconstruction filterbank for DNN-based sound source enhancement

Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi +2

We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to est…