3 papers
cs.SD2025
Dynamic Fusion Multimodal Network for SpeechWellness Detection
Wenqiang Sun, Han Yin, Jisheng Bai +1
Suicide is one of the leading causes of death among adolescents. Previous suicide risk prediction studies have primarily focused on either textual or acoustic information in isolat…
eess.AS2024
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
Jisheng Bai, Haohe Liu, Mou Wang +5
With the emergence of audio-language models, constructing large-scale paired audio-language datasets has become essential yet challenging for model development, primarily due to th…
eess.AS2024
Exploring Text-Queried Sound Event Detection with Audio Source Separation
Han Yin, Jisheng Bai, Yang Xiao +6
In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor…