5 papers
Grounded 3D-Aware Spatial Vision-Language Modeling
An-Chieh Cheng, Yang Fu, Yatai Ji +12
We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding-…
Inferring Dynamic Physical Properties from Video Foundation Models
Guanqi Zhan, Xianzheng Ma, Weidi Xie +1
We study the task of predicting dynamic physical properties from videos. More specifically, we consider physical properties that require temporal information to be inferred: elasti…
EGM: Efficient Visual Grounding Language Models
Guanqi Zhan, Changye Li, Zhijian Liu +4
Visual grounding is an essential capability of Visual Language Models (VLMs) to understand the real physical world. Previous state-of-the-art grounding visual language models usual…
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
Guanqi Zhan, Yuanpei Liu, Kai Han +2
The objective in this paper is to improve the performance of text-to-image retrieval. To this end, we introduce a new framework that can boost the performance of large-scale pre-tr…
Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under Occlusions
Ruihai Wu, Kai Cheng, Yan Shen +3
Perceiving and manipulating 3D articulated objects in diverse environments is essential for home-assistant robots. Recent studies have shown that point-level affordance provides ac…