4 papers
EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations
Yangcen Liu, Shuo Cheng, Xinchen Yin +6
Dexterous manipulation is limited by the cost of collecting large-scale robot demonstrations. Egocentric human videos offer a scalable source of diverse manipulation behaviors, but…
From Inference Efficiency to Embodied Efficiency: Revisiting Efficiency Metrics for Vision-Language-Action Models
Zhuofan Li, Hongkun Yang, Zhenyang Chen +4
Vision-Language-Action (VLA) models have recently enabled embodied agents to perform increasingly complex tasks by jointly reasoning over visual, linguistic, and motor modalities.…
ReSteer: Quantifying and Refining the Steerability of Multitask Robot Policies
Zhenyang Chen, Alan Tian, Liquan Wang +5
Despite strong multi-task pretraining, existing policies often exhibit poor task steerability. For example, a robot may fail to respond to a new instruction ``put the bowl in the s…
DemoGen: Synthetic Demonstration Generation for Data-Efficient Visuomotor Policy Learning
Zhengrong Xue, Shuying Deng, Zhenyang Chen +3
Visuomotor policies have shown great promise in robotic manipulation but often require substantial amounts of human-collected data for effective performance. A key reason underlyin…