works on

From the 2 of 130 linked papers with an AI index.

activity
20242026
most citedMindCube: Spatial Mental Modeling from Limited Views

1 citations · 1 across the 41 of their papers we have counts for

collaborators
Showing cs.ROShow all

27 papers · 1 filter

cs.RO2026

CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation

Letian Fu, Justin Yu, Karim El-Refai +13

"Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness as autonomous controllers for embodied ma…

cs.RO2026

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation

Haonan Chen, Yuxiang Ma, Stephen Tian +7

Long-horizon, contact-rich complex manipulation tasks, such as seating a GPU into a PCIe slot, demand both millimeter high precision and out-of-the-box generalization to new tasks.…

cs.RO2026

InSight: Self-Guided Skill Acquisition via Steerable VLAs

Maggie Wang, Lars Osterberg, Stephen Tian +3

Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. We present InSight, a…

cs.RO2026

DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?

Jadelynn Dao, Milan Ganai, Yasmina Abukhadra +7

Vision-Language Models (VLMs) are increasingly deployed as high-level planners for embodied agents, with an emerging strategy of scaling test-time compute to improve capability. Ho…

cs.RO2026

StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception

Evans Han, Yunfan Jiang, Yingke Wang +6

Recent advances in robot imitation learning have produced powerful visuomotor policies that manipulate diverse objects from visual inputs. However, monocular observations lack dept…

cs.RO2026

World Model for Robot Learning: A Comprehensive Survey

Bohan Hou, Gen Li, Jindou Jia +15

World models, which are predictive representations of how environments evolve under actions, have become a central component of robot learning. They support policy learning, planni…