activity
20242026
collaborators

9 papers

cs.MM2026

Learning to Prefer Reliably: Error-Augmented Emotion Preference Optimization with Calibrated Fusion

Zilong Huang, Junyi Peng, Junjie Li +5

Emotion preference learning uses pairwise comparisons between candidate descriptions to align multimodal large language models (MLLMs) with human judgments of open-ended emotion de…

cs.SD2026

Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning

Zezhong Jin, Xiaoyu Wang, Zhe Li +4

Residual connections are fundamental to deep speaker recogni- tion models, such as ECAPA-TDNN and ResNet. However, standard identity mapping limits information flow to a sin- gle p…

cs.MM2026

EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation

Zilong Huang, Kong Aik Lee, Junjie Li +2

Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC ofte…

cs.MM2026

EII-SCL: Harnessing Emotional Inertia for Multimodal Emotion Recognition in Conversation

Zilong Huang, Kong Aik Lee, Chong-Xin Gan +3

Multimodal emotion recognition in conversation (MERC) achieves accurate predictions by integrating multimodal and contextual information in dialogues. While current MERC approaches…

eess.AS2026

UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition

Chong-Xin Gan, Peter Bell, Man-Wai Mak +4

The joint training of speech enhancement and speaker embedding networks for speaker recognition is widely adopted under noisy acoustic environments. While effective, this paradigm…

cs.SD2026

Xi+: Uncertainty Supervision for Robust Speaker Embedding

Junjie Li, Kong Aik Lee, Duc-Tuan Truong +2

There are various factors that can influence the performance of speaker recognition systems, such as emotion, language and other speaker-related or context-related variations. Sinc…