collaborators

8 papers

eess.AS2025

Region-Specific Audio Tagging for Spatial Sound

Jinzheng Zhao, Yong Xu, Haohe Liu +6

Audio tagging aims to label sound events appearing in an audio recording. In this paper, we propose region-specific audio tagging, a new task which labels sound events in a given r…

eess.AS2024

Separate Anything You Describe

Xubo Liu, Qiuqiang Kong, Yan Zhao +7

Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given…

eess.AS2024

Universal Sound Separation with Self-Supervised Audio Masked Autoencoder

Junqi Zhao, Xubo Liu, Jinzheng Zhao +4

Universal sound separation (USS) is a task of separating mixtures of arbitrary sound sources. Typically, universal separation models are trained from scratch in a supervised manner…

eess.AS2024

Selective-Memory Meta-Learning with Environment Representations for Sound Event Localization and Detection

Jinbo Hu, Yin Cao, Ming Wu +4

Environment shifts and conflicts present significant challenges for learning-based sound event localization and detection (SELD) methods. SELD systems, when trained in particular a…

eess.AS2024

WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Xinhao Mei, Chutong Meng, Haohe Liu +6

The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years. However, researchers face challenges due to the costly and time-consuming col…

eess.AS2024

Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion

Zhichao Wang, Liumeng Xue, Qiuqiang Kong +4

Zero-shot voice conversion (VC) converts source speech into the voice of any desired speaker using only one utterance of the speaker without requiring additional model updates. Typ…