activity
20172022
most citedAudio Captioning Transformer

32 citations · 76 across the 6 of their papers we have counts for

collaborators

6 papers

eess.AS2022

Separate What You Describe: Language-Queried Audio Source Separation

Xubo Liu, Haohe Liu, Qiuqiang Kong +5

In this paper, we introduce the task of language-queried audio source separation (LASS), which aims to separate a target source from an audio mixture based on a natural language qu…

eess.AS2022

Leveraging Pre-trained BERT for Audio Captioning

Xubo Liu, Xinhao Mei, Qiushi Huang +6

Audio captioning aims at using natural language to describe the content of an audio clip. Existing audio captioning systems are generally based on an encoder-decoder architecture,…

eess.AS202120 cited

An Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning

Xinhao Mei, Qiushi Huang, Xubo Liu +10

Automated audio captioning aims to use natural language to describe the content of audio data. This paper presents an audio captioning system with an encoder-decoder architecture,…

eess.AS202132 cited

Audio Captioning Transformer

Xinhao Mei, Xubo Liu, Qiushi Huang +2

Audio captioning aims to automatically generate a natural language description of an audio clip. Most captioning models follow an encoder-decoder architecture, where the decoder pr…

eess.AS20212 cited

Conditional Sound Generation Using Neural Discrete Time-Frequency Representation Learning

Xubo Liu, Turab Iqbal, Jinzheng Zhao +3

Deep generative models have recently achieved impressive performance in speech and music synthesis. However, compared to the generation of those domain-specific sounds, generating…

cs.SI201722 cited

Sequential Prediction of Social Media Popularity with Deep Temporal Context Networks

Bo Wu, Wen-Huang Cheng, Yongdong Zhang +3

Prediction of popularity has profound impact for social media, since it offers opportunities to reveal individual preference and public attention from evolutionary social systems.…