activity
20242026
collaborators

9 papers

cs.CV2026

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer

Bohao Xing, Deng Li, Rong Gao +2

Recently, Transformer has made significant progress in various vision tasks. To balance computation and efficiency in video tasks, recent works heavily rely on factorized or window…

cs.CV2025

DEEMO: De-identity Multimodal Emotion Recognition and Reasoning

Deng Li, Bohao Xing, Xin Liu +3

Emotion understanding is a critical yet challenging task. Most existing approaches rely heavily on identity-sensitive information, such as facial expressions and speech, which rais…

cs.CV2025

MSF-Mamba: Motion-aware State Fusion Mamba for Efficient Micro-Gesture Recognition

Deng Li, Jun Shao, Bohao Xing +4

Micro-gesture recognition (MGR) targets the identification of subtle and fine-grained human motions and requires accurate modeling of both long-range and local spatiotemporal depen…

cs.CV2025

EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models

Bohao Xing, Xin Liu, Guoying Zhao +3

Emotion understanding is a critical yet challenging task. Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced their capabilities in this area. H…

cs.CV2025

FSBench: A Figure Skating Benchmark for Advancing Artistic Sports Understanding

Rong Gao, Xin Liu, Zhuozhao Hu +4

Figure skating, known as the "Art on Ice," is among the most artistic sports, challenging to understand due to its blend of technical elements (like jumps and spins) and overall ar…

cs.CV2025

AU-TTT: Vision Test-Time Training model for Facial Action Unit Detection

Bohao Xing, Kaishen Yuan, Zitong Yu +2

Facial Action Units (AUs) detection is a cornerstone of objective facial expression analysis and a critical focus in affective computing. Despite its importance, AU detection faces…