21 citations · 25 across the 6 of their papers we have counts for
6 papers
InViG: Benchmarking Interactive Visual Grounding with 500K Human-Robot Interactions
Hanbo Zhang, Jie Xu, Yuchen Mo +1
Ambiguity is ubiquitous in human communication. Previous approaches in Human-Robot Interaction (HRI) have often relied on predefined interaction templates, leading to reduced perfo…
Exploring Visual Pre-training for Robot Manipulation: Datasets, Models and Methods
Ya Jing, Xuelin Zhu, Xingbin Liu +4
Visual pre-training with large-scale real-world data has made great progress in recent years, showing great potential in robot learning with pixel observations. However, the recipe…
What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?
Yan Zeng, Hanbo Zhang, Jiani Zheng +5
Recent advancements in Large Language Models (LLMs) such as GPT4 have displayed exceptional multi-modal capabilities in following open-ended instructions given images. However, the…
ClickSeg: 3D Instance Segmentation with Click-Level Weak Annotations
Leyao Liu, Tao Kong, Minzhao Zhu +2
3D instance segmentation methods often require fully-annotated dense labels for training, which are costly to obtain. In this paper, we present ClickSeg, a novel click-level weakly…
Learning to Explore Informative Trajectories and Samples for Embodied Perception
Ya Jing, Tao Kong
We are witnessing significant progress on perception models, specifically those trained on large-scale internet images. However, efficiently generalizing these perception models to…
3D Part Assembly Generation with Instance Encoded Transformer
Rufeng Zhang, Tao Kong, Weihao Wang +2
It is desirable to enable robots capable of automatic assembly. Structural understanding of object parts plays a crucial role in this task yet remains relatively unexplored. In thi…