From the 1 of 14 linked papers with an AI index.
14 papers
HelloWorld: Enabling Socially Interactive Characters in Video World Models
Liangyang Ouyang, Ruicong Liu, Xuangeng Chu +2
Despite the remarkable recent progress of video world models, social interaction between users and the characters within these worlds remains unsupported. To fill this gap, we pres…
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
Gong Sitong, Tianyu Yan, Caixin Kang +6
The paper introduces Vinci2, a proactive on‑device assistant for continuous egocentric video that decides when to intervene by using memory‑augmented reasoning, and presents EgoSer…
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
Mingfang Zhang, Jingjing Pan, Ashutosh Kumar +7
Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of…
HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
Masatoshi Tateno, Gido Kato, Hirokatsu Kataoka +2
Hand-object interaction (HOI) inherently involves dynamics where human manipulations produce distinct spatio-temporal effects on objects. However, existing semantic HOI benchmarks…
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
Zhifan Zhu, Yifei Huang, Yoichi Sato +1
Humans can intuitively parallelise complex activities, but can a model predict this from observing a single person? Given one egocentric video, we introduce the N-Body Problem: pre…
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
Caixin Kang, Tianyu Yan, Sitong Gong +8
Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchmarks evaluate this capability…