32 citations · 76 across the 6 of their papers we have counts for
6 papers
Separate What You Describe: Language-Queried Audio Source Separation
Xubo Liu, Haohe Liu, Qiuqiang Kong +5
In this paper, we introduce the task of language-queried audio source separation (LASS), which aims to separate a target source from an audio mixture based on a natural language qu…
Leveraging Pre-trained BERT for Audio Captioning
Xubo Liu, Xinhao Mei, Qiushi Huang +6
Audio captioning aims at using natural language to describe the content of an audio clip. Existing audio captioning systems are generally based on an encoder-decoder architecture,…
An Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning
Xinhao Mei, Qiushi Huang, Xubo Liu +10
Automated audio captioning aims to use natural language to describe the content of audio data. This paper presents an audio captioning system with an encoder-decoder architecture,…
Audio Captioning Transformer
Xinhao Mei, Xubo Liu, Qiushi Huang +2
Audio captioning aims to automatically generate a natural language description of an audio clip. Most captioning models follow an encoder-decoder architecture, where the decoder pr…
Conditional Sound Generation Using Neural Discrete Time-Frequency Representation Learning
Xubo Liu, Turab Iqbal, Jinzheng Zhao +3
Deep generative models have recently achieved impressive performance in speech and music synthesis. However, compared to the generation of those domain-specific sounds, generating…
Sequential Prediction of Social Media Popularity with Deep Temporal Context Networks
Bo Wu, Wen-Huang Cheng, Yongdong Zhang +3
Prediction of popularity has profound impact for social media, since it offers opportunities to reveal individual preference and public attention from evolutionary social systems.…