44 citations · 44 across the 12 of their papers we have counts for
17 papers · 1 filter
ViTacPhys: Physical Property-Aware Grasping from Human Visual-Tactile Demonstrations
Yiwen Liu, Yujun Zhu, Kui Jia +3
Recent vision-based action models have demonstrated strong capabilities in complex manipulation, but they rarely leverage explicit object physical properties to adapt their policie…
Vid2WAM: Distilling Video Diffusion Priors into World Action Models
Chenhao Qiu, Ruixiang Wang, Runyi Zhao +7
World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generalization remain constrained by…
BiNoMaP: Learning Category-Level Bimanual Non-Prehensile Manipulation Primitives
Huayi Zhou, Kui Jia
Non-prehensile manipulation, encompassing ungraspable actions such as pushing, poking, pivoting, and wrapping, remains underexplored due to its contact-rich and analytically intrac…
Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning
Huayi Zhou, Wei Gao, Dekun Lu +12
End-to-end manipulation policies, combined with web-scale pretrained Vision-Language Models (VLMs), show the promise for generalizable and dexterous robotic manipulation. However,…
From Reaction to Anticipation: Proactive Failure Recovery through Agentic Task Graph for Robotic Manipulation
Sheng Xu, Ruixing Jin, Huayi Zhou +6
Although robotic manipulation has made significant progress, reliable execution remains challenging because task failures are inevitable in dynamic and unstructured environments. T…
VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
Huayi Zhou, Kui Jia
Achieving generalizable bimanual manipulation requires systems that can learn efficiently from minimal human input while adapting to real-world uncertainties and diverse embodiment…