6 papers
The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models
Xiangyu Zhang, Yuxin Li, Haoyang Zhang +5
The pursuit of a "unified" discrete token for both speech understanding and generation has led the Speech Language Model (SLM) community to heavily rely on Word Error Rate (WER) --…
Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective
Sijie Mai, Shiqin Han
Multimodal affective computing aims to predict humans' sentiment, emotion, intention, and opinion using language, acoustic, and visual modalities. However, current models often lea…
Addressing Missing and Noisy Modalities in One Solution: Unified Modality-Quality Framework for Low-quality Multimodal Data
Sijie Mai, Shiqin Han, Haifeng Hu
Multimodal data encountered in real-world scenarios are typically of low quality, with noisy modalities and missing modalities being typical forms that severely hinder model perfor…
CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion
Sijie Mai, Shiqin Han
Modality gap significantly restricts the effectiveness of multimodal fusion. Previous methods often use techniques such as diffusion models and adversarial learning to reduce the m…
Uncertainty-Aware Collaborative System of Large and Small Models for Multimodal Sentiment Analysis
Shiqin Han, Manning Gao, Menghua Jiang +3
Multimodal Large Language Models (MLLMs) have notably enhanced the performance of Multimodal Sentiment Analysis (MSA), yet their massive parameter scale leads to excessive resource…
GRCF: Two-Stage Groupwise Ranking and Calibration Framework for Multimodal Sentiment Analysis
Manning Gao, Leheng Zhang, Shiqin Han +3
Most Multimodal Sentiment Analysis research has focused on point-wise regression. While straightforward, this approach is sensitive to label noise and neglects whether one sample i…