activity
20172022
most citedCross-task learning for audio tagging, sound event detection and spatial localization: DCASE 2019 baseline systems

36 citations · 179 across the 24 of their papers we have counts for

collaborators
Showing eess.ASShow all

11 papers · 1 filter

eess.AS20222 cited

Multi-dimensional Edge-based Audio Event Relational Graph Representation Learning for Acoustic Scene Classification

Yuanbo Hou, Siyang Song, Chuang Yu +3

Most existing deep learning-based acoustic scene classification (ASC) approaches directly utilize representations extracted from spectrograms to identify target scenes. However, th…

eess.AS2022

Separate What You Describe: Language-Queried Audio Source Separation

Xubo Liu, Haohe Liu, Qiuqiang Kong +5

In this paper, we introduce the task of language-queried audio source separation (LASS), which aims to separate a target source from an audio mixture based on a natural language qu…

eess.AS2022

Leveraging Pre-trained BERT for Audio Captioning

Xubo Liu, Xinhao Mei, Qiushi Huang +6

Audio captioning aims at using natural language to describe the content of an audio clip. Existing audio captioning systems are generally based on an encoder-decoder architecture,…

eess.AS20222 cited

Deep Neural Decision Forest for Acoustic Scene Classification

Jianyuan Sun, Xubo Liu, Xinhao Mei +4

Acoustic scene classification (ASC) aims to classify an audio clip based on the characteristic of the recording environment. In this regard, deep learning based approaches have eme…

eess.AS202120 cited

An Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning

Xinhao Mei, Qiushi Huang, Xubo Liu +10

Automated audio captioning aims to use natural language to describe the content of audio data. This paper presents an audio captioning system with an encoder-decoder architecture,…

eess.AS202132 cited

Audio Captioning Transformer

Xinhao Mei, Xubo Liu, Qiushi Huang +2

Audio captioning aims to automatically generate a natural language description of an audio clip. Most captioning models follow an encoder-decoder architecture, where the decoder pr…