EmoVerse: Exploring Multimodal Large Language Models for Sentiment and Emotion Understanding
arXiv:2412.08049 · doi:10.1016/j.neucom.2025.130810
Abstract
Sentiment and emotion understanding are essential to applications such as human-computer interaction and depression detection. While Multimodal Large Language Models (MLLMs) demonstrate robust general capabilities, they face considerable challenges in the field of affective computing, particularly in detecting subtle facial expressions and handling complex emotion-related tasks, such as emotion reason inference and understanding emotions in long-context scenarios. Furthermore, there is a lack of a unified MLLM that can effectively handle both sentiment and emotion-related tasks. To address these challenges, we explore multi-task training strategies for MLLMs in affective computing and introduce Emotion Universe (EmoVerse), an MLLM designed to handle a broad spectrum of sentiment and emotion-related tasks. In addition, EmoVerse is capable of deeply analyzing the underlying causes of emotional states. We also introduce the Affective Multitask (AMT) Dataset, which supports multimodal sentiment analysis, multimodal emotion recognition, facial expression recognition, emotion reason inference, and emotion cause-pair extraction tasks. Extensive experiments demonstrate that EmoVerse outperforms existing methods, achieving state-of-the-art results in sentiment and emotion-related tasks. The code is available at https://github.com/liaolea/EmoVerse.
References in corpus (10)
- LoRA: Low-Rank Adaptation of Large Language Models
- Visual Instruction Tuning
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- GA2MIF: Graph and Attention Based Two-Stage Multi-Source Information Fusion for Conversational Emotion Detection
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
- InstructERC: Reforming Emotion Recognition in Conversation with Multi-task Retrieval-Augmented Large Language Models
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
- TimeCMA: Towards LLM-Empowered Multivariate Time Series Forecasting via Cross-Modality Alignment
- LLaVA-Video: Video Instruction Tuning With Synthetic Data