12 papers
PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models
Zihan Song, Shuo Ye, Bo Zhao +4
Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindranc…
GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs
Taorui Wang, Wei Xia, Hui Ma +5
Micro-gesture recognition demands the detection of fleeting, spatially localized movements that are frequently overwhelmed by dominant static appearances and background noise. Whil…
AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
Jiayu Zhang, Shuo Ye, Qilang Ye +3
Audio-Visual Question Answering (AVQA) requires models to effectively utilize both visual and auditory modalities to answer complex and diverse questions about audio-visual scenes.…
DeceptionX: From Multimodal Evidence to Explainable Deception Detection
Jiayu Zhang, Shuo Ye, Jiajian Huang +8
Deception detection is a critical and highly challenging task within affective computing and behavioral analysis. Existing deep learning methods typically treat this task as a stra…
Viewport-Unaware Blind Omnidirectional Image Quality Assessment: A Unified and Generalized Approach
Jiebin Yan, Kangcheng Wu, Jingwen Hou +3
Blind omnidirectional image quality assessment (BOIQA) presents a great challenge to the visual quality assessment community, due to different storage formats and diverse user view…
AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition
Zeheng Wang, Zitong Yu, Yijie Zhu +9
LLM-based multimodal emotion recognition relies on static parametric memory and often hallucinates when interpreting nuanced affective states. In this paper, given that single-roun…