Publications (8)
I Know What You Draw: Learning Grasp Detection Conditioned on a Few Freehand Sketches
Haitao Lin, Chilam Cheang, Yanwei Fu +1
In this paper, we are interested in the problem of generating target grasps by understanding freehand sketches. The sketch is useful for the persons who cannot formulate language a…
SAR-Net: Shape Alignment and Recovery Network for Category-level 6D Object Pose and Size Estimation
Haitao Lin, Zichang Liu, Chilam Cheang +3
Given a single scene image, this paper proposes a method of Category-level 6D Object Pose and Size Estimation (COPSE) from the point cloud of the target object, without external re…
Vision-Language Foundation Models as Effective Robot Imitators
Xinghang Li, Minghuan Liu, Hanbo Zhang +9
Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipul…
IRASim: A Fine-Grained World Model for Robot Manipulation
Fangqi Zhu, Hongtao Wu, Song Guo +3
World models allow autonomous agents to plan and explore by predicting the visual outcomes of different actions. However, for robot manipulation, it is challenging to accurately mo…
GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy
Peiyan Li, Hongtao Wu, Yan Huang +3
The robotics community has consistently aimed to achieve generalizable robot manipulation with flexible natural language instructions. One primary challenge is that obtaining robot…
Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation
Hongtao Wu, Ya Jing, Chilam Cheang +6
Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of th…