From the 1 of 7 linked papers with an AI index.
7 papers
MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos
Jinpeng Hu, Erqiang Wang, Shan Wang +4
The paper presents MMHBench, a multimodal benchmark of 268 long-form videos with 2,184 questions designed to evaluate mental health understanding from both observable behavior and…
Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset
Zhiyan Zhang, Peipei Song, Jinpeng Hu +3
Video emotion analysis is typically framed as a static classification problem, treating each clip as an independent labeled unit. However, such a formulation overlooks a key psycho…
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning
Zhiyuan Han, Beier Zhu, Wenwen Tong +6
We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit u…
A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation
Liping Wang, Cheng Ye, Weidong Chen +3
Multimodal empathetic response generation (MERG) aims to generate emotionally engaging and empathetic responses based on users' multimodal contexts. Existing approaches usually rel…
FACE-net: Factual Calibration and Emotion Augmentation for Retrieval-enhanced Emotional Video Captioning
Weidong Chen, Cheng Ye, Zhendong Mao +5
Emotional Video Captioning (EVC) is an emerging task, which aims to describe factual content with the intrinsic emotions expressed in videos. Existing works perceive global emotion…
Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning
Zhiyuan Han, Beier Zhu, Yanlong Xu +2
Despite their strong performance in multimodal emotion reasoning, existing Multimodal Large Language Models (MLLMs) often overlook the scenarios involving emotion conflicts, where…