6 papers
Mixture of Horizons in Action Chunking
Dong Jing, Gang Wang, Jiaqi Liu +7
Vision-language-action (VLA) models have shown remarkable capabilities in robotic manipulation, but their performance is sensitive to the used during…
Rethinking Intermediate Representation for VLM-based Robot Manipulation
Weiliang Tang, Jialin Gao, Jia-Hui Pan +6
Vision-Language Model (VLM) is an important component to enable robust robot manipulation. Yet, using it to translate human instructions into an action-resolvable intermediate repr…
OPA-Pack: Object-Property-Aware Robotic Bin Packing
Jia-Hui Pan, Yeok Tatt Cheah, Zhengzhe Liu +5
Robotic bin packing aids in a wide range of real-world scenarios such as e-commerce and warehouses. Yet, existing works focus mainly on considering the shape of objects to optimize…
Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation
Weiliang Tang, Dong Jing, Jia-Hui Pan +5
Recent Large Multimodal Models have demonstrated remarkable reasoning capabilities, especially in solving complex mathematical problems and realizing accurate spatial perception. O…
Overcoming Support Dilution for Robust Few-shot Semantic Segmentation
Wailing Tang, Biqi Yang, Pheng-Ann Heng +2
Few-shot Semantic Segmentation (FSS) is a challenging task that utilizes limited support images to segment associated unseen objects in query images. However, recent FSS methods ar…
GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation
Weiliang Tang, Jia-Hui Pan, Yun-Hui Liu +4
We present GeoManip, a framework to enable generalist robots to leverage essential conditions derived from object and part relationships, as geometric constraints, for robot manipu…