7 papers
EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos
Yifan Zhong, Zhang Chen, Tianrui Guan +13
Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accur…
System Design for Maintaining Internal State Consistency in Long-Horizon Robotic Tabletop Games
Guangyu Zhao, Ceyao Zhang, Chengdong Ma +16
Long-horizon tabletop games pose a distinct systems challenge for robotics: small perceptual or execution errors can invalidate accumulated task state, propagate across decision-ma…
DexKnot: Generalizable Visuomotor Policy Learning for Dexterous Bag-Knotting Manipulation
Jiayuan Zhang, Ruihai Wu, Haojun Chen +5
Knotting plastic bags is a common task in daily life, yet it is challenging for robots due to the bags' infinite degrees of freedom and complex physical dynamics. Existing methods…
DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping
Yifan Zhong, Xuchuan Huang, Ruochong Li +9
Dexterous grasping remains a fundamental yet challenging problem in robotics. A general-purpose robot must be capable of grasping diverse objects in arbitrary scenarios. However, e…
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
Yifan Zhong, Fengshuo Bai, Shaofei Cai +11
The remarkable advancements of vision and language foundation models in multimodal understanding, reasoning, and generation has sparked growing efforts to extend such intelligence…
Falcon: Fast Visuomotor Policies via Partial Denoising
Haojun Chen, Minghao Liu, Chengdong Ma +8
Diffusion policies are widely adopted in complex visuomotor tasks for their ability to capture multimodal action distributions. However, the multiple sampling steps required for ac…