works on

From the 1 of 17 linked papers with an AI index.

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert

Mingyu Liu, Zheng Huang, Xiaoyi Lin +6

Vision-language models demonstrate strong reasoning and planning abilities, yet grounding these predictions into precise robot actions remains a central challenge. Existing Vision-…

cs.CV2026

ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning

Muzhi Zhu, Hao Zhong, Canyu Zhao +9

Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information. It is a critical component of effic…

cs.CV2026

AGILE: Hand-Object Interaction Reconstruction from Video via Agentic Generation

Jin-Chuan Shi, Binhong Ye, Tao Liu +6

Reconstructing dynamic hand-object interactions from monocular videos is critical for dexterous manipulation data collection and creating realistic digital twins for robotics and V…

cs.CV2026

RegionReasoner: Region-Grounded Multi-Round Visual Reasoning

Wenfang Sun, Hao Chen, Yingjun Du +2

Large vision-language models have achieved remarkable progress in visual reasoning, yet most existing systems rely on single-step or text-only reasoning, limiting their ability to…

cs.CV2026

Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video

Zihui Gao, Ke Liu, Donny Y. Chen +4

Geometric foundation models show promise in 3D reconstruction, yet their progress is severely constrained by the scarcity of diverse, large-scale 3D annotations. While Internet vid…

cs.CV2026

CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning

Wenxin Ma, Chenlong Wang, Ruisheng Yuan +6

Humans can look at a static scene and instantly predict what happens next -- will moving this object cause a collision? We call this ability Causal Spatial Reasoning. However, curr…