24 papers
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
Mingfang Zhang, Jingjing Pan, Ashutosh Kumar +7
Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of…
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
Zhifan Zhu, Yifei Huang, Yoichi Sato +1
Humans can intuitively parallelise complex activities, but can a model predict this from observing a single person? Given one egocentric video, we introduce the N-Body Problem: pre…
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
Caixin Kang, Tianyu Yan, Sitong Gong +8
Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchmarks evaluate this capability…
EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning
Zeyu Wang, Chang Liu, Eduardus Tjitrahardja +22
Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI assistant experiences, remain…
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting
Ruicong Liu, Yifei Huang, Liangyang Ouyang +2
Real-time 3D hand forecasting is a critical component for fluid human-computer interaction in applications like AR and assistive robotics. However, existing methods are ill-suited…
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
Jiacheng Hua, Yishu Yin, Yuhang Wu +3
Existing Multimodal Large Language Models (MLLMs) struggle with 3D spatial reasoning, as they fail to construct structured abstractions of the 3D environment depicted in video inpu…