8 papers
MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement
Daiqing Wu, Dongbao Yang, Jiashu Yao +4
Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intellige…
Multimodal Emotion Recognition with Large Language Models
Hongrui Zhang, Daiqing Wu, Yangyang Li +4
Multimodal Emotion Recognition (MER) focuses on identifying and interpreting emotions from modality-compound inputs. Closely mirroring human cognitive processes in real-world envir…
AffectVerse: Emotional World Models for Multimodal Affective Computing
Bo Zhao, Fanghua Ye, Yixin Ji +3
Humans infer emotions by integrating observed multimodal cues with expectations about how affective states may unfold. Existing multimodal large language models (MLLMs), however, o…
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
Zihan Tang, Leqi Shen, Hui Chen +7
Vision-Language Models (VLMs) have shown strong promise on Optical Character Recognition (OCR), yet the sheer number of visual tokens required to encode dense documents incurs proh…
HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
Tingting Han, Xinsong Tao, Yufei Yin +3
Temporal Sentence Grounding in Videos (TSGV) aims to temporally localize segments of a video that correspond to a given natural language query. Despite recent progress, most existi…
Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning
Daiqing Wu, Xuan Zhang, Dongbao Yang +7
The maturation of Large Audio Language Models (LALMs) has raised growing expectations for them to comprehend complex audio much like humans. Current efforts primarily replicate tex…