6 papers
DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
Yuhua Wen, Qifei Li, Yingying Zhou +4
Multimodal sentiment analysis (MSA) integrates various modalities, such as text, image, and audio, to provide a more comprehensive understanding of sentiment. However, effective MS…
HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
Bingsong Bai, Yizhong Geng, Fengping Wang +4
Zero-shot singing voice conversion (SVC) transforms a source singer's timbre to an unseen target speaker's voice while preserving melodic content without fine-tuning. Existing meth…
SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding
Bingsong Bai, Qihang Lu, Wenbing Yang +8
Paralinguistic sounds, like laughter and sighs, are crucial for synthesizing more realistic and engaging speech. However, existing methods typically depend on proprietary datasets,…
Deep Learning Approaches for Multimodal Intent Recognition: A Survey
Jingwei Zhao, Yuhua Wen, Qifei Li +8
Intent recognition aims to identify users' underlying intentions, traditionally focusing on text in natural language processing. With growing demands for natural human-computer int…
Psy-Copilot: Visual Chain of Thought for Counseling
Keqi Chen, Zekai Sun, Huijun Lian +2
Large language models (LLMs) are becoming increasingly popular in the field of psychological counseling. However, when human therapists work with LLMs in therapy sessions, it is ha…
Psy-Insight: Explainable Multi-turn Bilingual Dataset for Mental Health Counseling
Keqi Chen, Zekai Sun, Yuhua Wen +3
The in-context learning capabilities of large language models (LLMs) show great potential in mental health support. However, the lack of counseling datasets, particularly in Chines…