8 papers
Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation Learning
Yiyao Ma, Kai Chen, Zhongxiang Zhou +5
Monocular 3D shape recovery is fundamental to geometric understanding, yet achieving robust generalization across arbitrary viewpoints and unseen object categories remains a signif…
Toward Embodiment Equivariant Vision-Language-Action Policy
Anzhe Chen, Yifei Yang, Zhenjie Zhu +4
Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel…
ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
Zhichen Lou, Kechun Xu, Zhongxiang Zhou +1
The advancement of embodied intelligence is accelerating the integration of robots into daily life as human assistants. This evolution requires robots to not only interpret high-le…
TOP: Time Optimization Policy for Stable and Accurate Standing Manipulation with Humanoid Robots
Zhenghan Chen, Haocheng Xu, Haodong Zhang +7
Humanoid robots have the potential capability to perform a diverse range of manipulation tasks, but this is based on a robust and precise standing controller. Existing methods are…
Disambiguate Gripper State in Grasp-Based Tasks: Pseudo-Tactile as Feedback Enables Pure Simulation Learning
Yifei Yang, Lu Chen, Zherui Song +5
Grasp-based manipulation tasks are fundamental to robots interacting with their environments, yet gripper state ambiguity significantly reduces the robustness of imitation learning…
CNSv2: Probabilistic Correspondence Encoded Neural Image Servo
Anzhe Chen, Hongxiang Yu, Shuxin Li +5
Visual servo based on traditional image matching methods often requires accurate keypoint correspondence for high precision control. However, keypoint detection or matching tends t…