papers
Publications (3)
cs.CL2024
Mutual Information-based Representations Disentanglement for Unaligned Multimodal Language Sequences
Fan Qian, Jiqing Han, Jianchen Li +3
The key challenge in unaligned multimodal language sequences lies in effectively integrating information from various modalities to obtain a refined multimodal joint representation…
cs.CL2026
LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models Using in-the-wild Data
Wen Ding, Fan Qian
Although state-of-the-art Speech Foundational Models can produce high-quality text pseudo-labels, applying Semi-Supervised Learning (SSL) for in-the-wild real-world data remains ch…
cs.SD2022
Contrastive Regularization for Multimodal Emotion Recognition Using Audio and Text
Fan Qian, Jiqing Han
Speech emotion recognition is a challenge and an important step towards more natural human-computer interaction (HCI). The popular approach is multimodal emotion recognition based…