2 papers
cs.CL2026
LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models Using in-the-wild Data
Wen Ding, Fan Qian
Although state-of-the-art Speech Foundational Models can produce high-quality text pseudo-labels, applying Semi-Supervised Learning (SSL) for in-the-wild real-world data remains ch…
cs.CL2024
Mutual Information-based Representations Disentanglement for Unaligned Multimodal Language Sequences
Fan Qian, Jiqing Han, Jianchen Li +3
The key challenge in unaligned multimodal language sequences lies in effectively integrating information from various modalities to obtain a refined multimodal joint representation…