6 papers
MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis
Yuhua Wen, Yingying Zhou, Qifei Li +4
Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounter incomplete or corrupted mo…
MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation
Yizhong Geng, Wenxin Fu, Kecan Mao +7
Neural audio codecs serve as fundamental tokenizers for LLM-based audio generation. While semantic priors are widely exploited to enhance linguistic intelligibility, the integratio…
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
Cong Wang, Yizhong Geng, Yuhua Wen +7
Speech emotion recognition (SER) is an important technology in human-computer interaction. However, achieving high performance is challenging due to emotional complexity and scarce…
DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
Yuhua Wen, Qifei Li, Yingying Zhou +4
Multimodal sentiment analysis (MSA) integrates various modalities, such as text, image, and audio, to provide a more comprehensive understanding of sentiment. However, effective MS…
Deep Learning Approaches for Multimodal Intent Recognition: A Survey
Jingwei Zhao, Yuhua Wen, Qifei Li +8
Intent recognition aims to identify users' underlying intentions, traditionally focusing on text in natural language processing. With growing demands for natural human-computer int…
LoD: Loss-difference OOD Detection by Intentionally Label-Noisifying Unlabeled Wild Data
Chuanxing Geng, Qifei Li, Xinrui Wang +3
Using unlabeled wild data containing both in-distribution (ID) and out-of-distribution (OOD) data to improve the safety and reliability of models has recently received increasing a…