38 papers
MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving
Zhijing Cheng, Xuancheng Zhang, Donglin Di +4
End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downs…
TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues
Junhui She, Fei Wang, Kun Li +4
Gaze target estimation aims to infer the position of a person's gaze within a scene. Within mainstream design logic, multi-branch methods require extra supervision and annotations,…
SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection
Wenlin Wu, Sheng Zhou, Peipei Song +3
As video generation paradigms evolve from localized manipulation to full-scene synthesis, AI-generated video detection becomes increasingly challenging, as forgeries exhibit cohere…
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy
Zhiyuan Han, Beier Zhu, Wenwen Tong +8
We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, even though it makes predictions more interpretable. Speci…
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning
Zhiyuan Han, Beier Zhu, Wenwen Tong +6
We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit u…
A New Multi-Domain Benchmark for Micro-Action Recognition and Detection
Yanbin Hao, Pengyu Liu, Xing Wei +3
Micro-actions are short-duration, low-amplitude subtle body movements at the whole-body level that can reveal latent intentions, involuntary reactions, and fine-grained affective c…