collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

ACE-Ego-Hand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

Yufei Liu, Xixi Wang, Hao Li +8

Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-o…

cs.CV2026

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

Weiliang Chen, Haowen Sun, Jun Gao +40

A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, w…

cs.CV2026

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

Yalun Dai, Hao Li, Shulin Tian +8

Real-world spatial intelligence requires reasoning over a continuous and evolving 3D world, yet existing VLMs and tool-augmented agents largely remain tied to static, stateless inf…

cs.CV2026

Apple-: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

Runmao Yao, Kairui Hu, Yukang Cao +11

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausi…

cs.CV2025

Unified Multimodal Models as Auto-Encoders

Zhiyuan Yan, Kaiqing Lin, Zongjian Li +10

Image-to-text (I2T) understanding and text-to-image (T2I) generation are two fundamental, important yet traditionally isolated multimodal tasks. Despite their intrinsic connection,…