papers

Publications (8)

cs.RO2022

I Know What You Draw: Learning Grasp Detection Conditioned on a Few Freehand Sketches

Haitao Lin, Chilam Cheang, Yanwei Fu +1

In this paper, we are interested in the problem of generating target grasps by understanding freehand sketches. The sketch is useful for the persons who cannot formulate language a…

cs.CV2022

SAR-Net: Shape Alignment and Recovery Network for Category-level 6D Object Pose and Size Estimation

Haitao Lin, Zichang Liu, Chilam Cheang +3

Given a single scene image, this paper proposes a method of Category-level 6D Object Pose and Size Estimation (COPSE) from the point cloud of the target object, without external re…

cs.RO2024

Vision-Language Foundation Models as Effective Robot Imitators

Xinghang Li, Minghuan Liu, Hanbo Zhang +9

Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipul…

cs.RO2025

IRASim: A Fine-Grained World Model for Robot Manipulation

Fangqi Zhu, Hongtao Wu, Song Guo +3

World models allow autonomous agents to plan and explore by predicting the visual outcomes of different actions. However, for robot manipulation, it is challenging to accurately mo…

cs.RO2024

GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy

Peiyan Li, Hongtao Wu, Yan Huang +3

The robotics community has consistently aimed to achieve generalizable robot manipulation with flexible natural language instructions. One primary challenge is that obtaining robot…

cs.RO2023

Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Hongtao Wu, Ya Jing, Chilam Cheang +6

Generative pre-trained models have demonstrated remarkable effectiveness in language and vision domains by learning useful representations. In this paper, we extend the scope of th…