activity
20182022
most citedThe NTT DCASE2020 Challenge Task 6 system: Automated Audio Captioning with Keywords and Sentence Length Estimation

21 citations · 34 across the 6 of their papers we have counts for

collaborators

10 papers

cs.CV20221 cited

Reflectance-Oriented Probabilistic Equalization for Image Enhancement

Xiaomeng Wu, Yongqing Sun, Akisato Kimura +1

Despite recent advances in image enhancement, it remains difficult for existing approaches to adaptively improve the brightness and contrast for both low-light and normal-light ima…

cs.CV2022

Reflectance-Guided, Contrast-Accumulated Histogram Equalization

Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino

Existing image enhancement methods fall short of expectations because with them it is difficult to improve global and local image contrast simultaneously. To address this problem,…

eess.AS2022

Composing General Audio Representation by Fusing Multilayer Features of a Pre-trained Model

Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2

Many application studies rely on audio DNN models pre-trained on a large-scale dataset as essential feature extractors, and they extract features from the last layers. In this stud…

eess.AS20211 cited

BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation

Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi +2

Inspired by the recent progress in self-supervised learning for computer vision that generates supervision using data augmentations, we explore a new general-purpose audio represen…

cs.CV2021

Attention to Warp: Deep Metric Learning for Multivariate Time Series

Shinnosuke Matsuo, Xiaomeng Wu, Gantugs Atarsaikhan +4

Deep time series metric learning is challenging due to the difficult trade-off between temporal invariance to nonlinear distortion and discriminative power in identifying non-match…

eess.AS202011 cited

Effects of Word-frequency based Pre- and Post- Processings for Audio Captioning

Daiki Takeuchi, Yuma Koizumi, Yasunori Ohishi +2

The system we used for Task 6 (Automated Audio Captioning)of the Detection and Classification of Acoustic Scenes and Events(DCASE) 2020 Challenge combines three elements, namely, d…