32 citations · 58 across the 7 of their papers we have counts for
7 papers
Automated Audio Captioning via Fusion of Low- and High- Dimensional Features
Jianyuan Sun, Xubo Liu, Xinhao Mei +3
Automated audio captioning (AAC) aims to describe the content of an audio clip using simple sentences. Existing AAC methods are developed based on an encoder-decoder architecture t…
Separate What You Describe: Language-Queried Audio Source Separation
Xubo Liu, Haohe Liu, Qiuqiang Kong +5
In this paper, we introduce the task of language-queried audio source separation (LASS), which aims to separate a target source from an audio mixture based on a natural language qu…
Leveraging Pre-trained BERT for Audio Captioning
Xubo Liu, Xinhao Mei, Qiushi Huang +6
Audio captioning aims at using natural language to describe the content of an audio clip. Existing audio captioning systems are generally based on an encoder-decoder architecture,…
Deep Neural Decision Forest for Acoustic Scene Classification
Jianyuan Sun, Xubo Liu, Xinhao Mei +4
Acoustic scene classification (ASC) aims to classify an audio clip based on the characteristic of the recording environment. In this regard, deep learning based approaches have eme…
An Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning
Xinhao Mei, Qiushi Huang, Xubo Liu +10
Automated audio captioning aims to use natural language to describe the content of audio data. This paper presents an audio captioning system with an encoder-decoder architecture,…
Audio Captioning Transformer
Xinhao Mei, Xubo Liu, Qiushi Huang +2
Audio captioning aims to automatically generate a natural language description of an audio clip. Most captioning models follow an encoder-decoder architecture, where the decoder pr…