3 papers
cs.CV2025
Point What You Mean: Visually Grounded Instruction Policy
Hang Yu, Juntu Zhao, Yufeng Liu +9
Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especial…
cs.RO2025
Best of Sim and Real: Decoupled Visuomotor Manipulation via Learning Control in Simulation and Perception in Real
Jialei Huang, Zhaoheng Yin, Yingdong Hu +3
Sim-to-real transfer remains a fundamental challenge in robot manipulation due to the entanglement of perception and control in end-to-end learning. We present a decoupled framewor…
cs.RO2025
Do You Need Proprioceptive States in Visuomotor Policies?
Juntu Zhao, Wenbo Lu, Di Zhang +10
Imitation-learning-based visuomotor policies have been widely used in robot manipulation, where both visual observations and proprioceptive states are typically adopted together fo…