13 papers
Enhancing Part-Level Point Grounding for Any Open-Source MLLMs
Jin-Cheng Jhang, Fu-En Wang, Xin Yang +4
Visual grounding aims to associate free-form textual queries with specific regions in an image. While recent Multimodal Large Language Models (MLLMs) have demonstrated promising ca…
Revisiting Model Stitching In the Foundation Model Era
Zheda Mai, Ke Zhang, Fu-En Wang +6
Model stitching, connecting early layers of one model (source) to later layers of another (target) via a light stitch layer, has served as a probe of representational compatibility…
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
Yurou Yang, Muyuan Lin, Roberto Martin-Martin +4
Recent work explores new opportunities at the intersection of vision-language-action models (VLAs) and geometric foundation models (GFMs) for 3D reconstruction, such as VGGT. While…
Explicit Memory through Online 3D Gaussian Splatting Improves Class-Agnostic Video Segmentation
Anthony Opipari, Aravindhan K Krishnan, Shreekant Gayaka +4
Remembering where object segments were predicted in the past is useful for improving the accuracy and consistency of class-agnostic video segmentation algorithms. Existing video se…
Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning
Houjian Yu, Zheming Zhou, Min Sun +5
Enabling robots to grasp objects specified through natural language is essential for effective human-robot interaction, yet it remains a significant challenge. Existing approaches…
OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations
Peng-Hao Hsu, Ke Zhang, Fu-En Wang +6
Open-vocabulary (OV) 3D object detection is an emerging field, yet its exploration through image-based methods remains limited compared to 3D point cloud-based methods. We introduc…