most citedAudio Captioning Transformer

32 citations · 56 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD20222 cited

Automated Audio Captioning via Fusion of Low- and High- Dimensional Features

Jianyuan Sun, Xubo Liu, Xinhao Mei +3

Automated audio captioning (AAC) aims to describe the content of an audio clip using simple sentences. Existing AAC methods are developed based on an encoder-decoder architecture t…

eess.AS2022

Separate What You Describe: Language-Queried Audio Source Separation

Xubo Liu, Haohe Liu, Qiuqiang Kong +5

In this paper, we introduce the task of language-queried audio source separation (LASS), which aims to separate a target source from an audio mixture based on a natural language qu…

eess.AS2022

Leveraging Pre-trained BERT for Audio Captioning

Xubo Liu, Xinhao Mei, Qiushi Huang +6

Audio captioning aims at using natural language to describe the content of an audio clip. Existing audio captioning systems are generally based on an encoder-decoder architecture,…

eess.AS20222 cited

Deep Neural Decision Forest for Acoustic Scene Classification

Jianyuan Sun, Xubo Liu, Xinhao Mei +4

Acoustic scene classification (ASC) aims to classify an audio clip based on the characteristic of the recording environment. In this regard, deep learning based approaches have eme…

eess.AS202120 cited

An Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning

Xinhao Mei, Qiushi Huang, Xubo Liu +10

Automated audio captioning aims to use natural language to describe the content of audio data. This paper presents an audio captioning system with an encoder-decoder architecture,…

eess.AS202132 cited

Audio Captioning Transformer

Xinhao Mei, Xubo Liu, Qiushi Huang +2

Audio captioning aims to automatically generate a natural language description of an audio clip. Most captioning models follow an encoder-decoder architecture, where the decoder pr…