6 papers
Choose What to Observe: Task-Aware Semantic-Geometric Representations for Visuomotor Policy
Haoran Ding, Liang Ma, Yaxun Yang +7
Visuomotor policies learned from demonstrations often overfit to nuisance visual factors in raw RGB observations, resulting in brittle behavior under appearance shifts such as back…
Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics
Open-H-Embodiment Consortium, :, Nigel Nelson +213
Autonomous medical robots hold promise to improve patient outcomes, reduce provider workload, democratize access to care, and enable superhuman precision. However, autonomous medic…
World2Act: Latent Action Post-Training from World Model Dynamics
An Dinh Vuong, Tuan Van Vo, Abdullah Sohail +6
World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve generalization under task and scene…
3D-CovDiffusion: 3D-Aware Diffusion Policy for Coverage Path Planning
Chenyuan Chen, Haoran Ding, Ran Ding +6
Diffusion models have shown strong potential for robot skill learning, yet their role in coverage path planning remains underexplored. In industrial surface processing (painting, p…
Imagination at Inference: Synthesizing In-Hand Views for Robust Visuomotor Policy Inference
Haoran Ding, Anqing Duan, Zezhou Sun +2
Visual observations from different viewpoints can significantly influence the performance of visuomotor policies in robotic manipulation. Among these, egocentric (in-hand) views of…
Towards Safe Imitation Learning via Potential Field-Guided Flow Matching
Haoran Ding, Anqing Duan, Zezhou Sun +4
Deep generative models, particularly diffusion and flow matching models, have recently shown remarkable potential in learning complex policies through imitation learning. However,…