5 papers
Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension
Yik Lung Pang, Changjae Oh
Given a textual description, the task of referring expression comprehension (REC) involves the localisation of the referred object in an image. Multimodal large language models (ML…
LaVA-Man: Learning Visual Action Representations for Robot Manipulation
Chaoran Zhu, Hengyi Wang, Yik Lung Pang +1
Visual-textual understanding is essential for language-guided robot manipulation. Recent works leverage pre-trained vision-language models to measure the similarity between encoded…
Toward Human-Robot Teaming: Learning Handover Behaviors from 3D Scenes
Yuekun Wu, Yik Lung Pang, Andrea Cavallaro +1
Human-robot teaming (HRT) systems often rely on large-scale datasets of human and robot interactions, especially for close-proximity collaboration tasks such as human-robot handove…
Learning human-to-robot handovers through 3D scene reconstruction
Yuekun Wu, Yik Lung Pang, Andrea Cavallaro +1
Learning robot manipulation policies from raw, real-world image data requires a large number of robot-action trials in the physical environment. Although training using simulations…
Stereo Hand-Object Reconstruction for Human-to-Robot Handover
Yik Lung Pang, Alessio Xompero, Changjae Oh +1
Jointly estimating hand and object shape facilitates the grasping task in human-to-robot handovers. However, relying on hand-crafted prior knowledge about the geometric structure o…