5 citations · 5 across the 4 of their papers we have counts for
5 papers · 1 filter
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
Jisheng Bai, Haohe Liu, Mou Wang +5
With the emergence of audio-language models, constructing large-scale paired audio-language datasets has become essential yet challenging for model development, primarily due to th…
Description on IEEE ICME 2024 Grand Challenge: Semi-supervised Acoustic Scene Classification under Domain Shift
Jisheng Bai, Mou Wang, Haohe Liu +11
Acoustic scene classification (ASC) is a crucial research problem in computational auditory scene analysis, and it aims to recognize the unique acoustic characteristics of an envir…
Sub-band and Full-band Interactive U-Net with DPRNN for Demixing Cross-talk Stereo Music
Han Yin, Mou Wang, Jisheng Bai +3
This paper presents a detailed description of our proposed methods for the ICASSP 2024 Cadenza Challenge. Experimental results show that the proposed system can achieve better perf…
Interactive Dual-Conformer with Scene-Inspired Mask for Soft Sound Event Detection
Han Yin, Jisheng Bai, Mou Wang +3
Traditional binary hard labels for sound event detection (SED) lack details about the complexity and variability of sound event distributions. Recently, a novel annotation workflow…
AudioLog: LLMs-Powered Long Audio Logging with Hybrid Token-Semantic Contrastive Learning
Jisheng Bai, Han Yin, Mou Wang +4
Previous studies in automated audio captioning have faced difficulties in accurately capturing the complete temporal details of acoustic scenes and events within long audio sequenc…