5 citations · 7 across the 7 of their papers we have counts for
7 papers
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
Jisheng Bai, Haohe Liu, Mou Wang +5
With the emergence of audio-language models, constructing large-scale paired audio-language datasets has become essential yet challenging for model development, primarily due to th…
Leveraging LLM and Text-Queried Separation for Noise-Robust Sound Event Detection
Han Yin, Yang Xiao, Jisheng Bai +1
Sound Event Detection (SED) is challenging in noisy environments where overlapping sounds obscure target events. Language-queried audio source separation (LASS) aims to isolate the…
Squeeze-and-Excite ResNet-Conformers for Sound Event Localization, Detection, and Distance Estimation for DCASE 2024 Challenge
Jun Wei Yeow, Ee-Leng Tan, Jisheng Bai +2
This technical report details our systems submitted for Task 3 of the DCASE 2024 Challenge: Audio and Audiovisual Sound Event Localization and Detection (SELD) with Source Distance…
FMSG-JLESS Submission for DCASE 2024 Task4 on Sound Event Detection with Heterogeneous Training Dataset and Potentially Missing Labels
Yang Xiao, Han Yin, Jisheng Bai +1
This report presents the systems developed and submitted by Fortemedia Singapore (FMSG) and Joint Laboratory of Environmental Sound Sensing (JLESS) for DCASE 2024 Task 4. The task…
Description on IEEE ICME 2024 Grand Challenge: Semi-supervised Acoustic Scene Classification under Domain Shift
Jisheng Bai, Mou Wang, Haohe Liu +11
Acoustic scene classification (ASC) is a crucial research problem in computational auditory scene analysis, and it aims to recognize the unique acoustic characteristics of an envir…
Sub-band and Full-band Interactive U-Net with DPRNN for Demixing Cross-talk Stereo Music
Han Yin, Mou Wang, Jisheng Bai +3
This paper presents a detailed description of our proposed methods for the ICASSP 2024 Cadenza Challenge. Experimental results show that the proposed system can achieve better perf…