activity
20242026
collaborators

13 papers

cs.CV2026

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs

Jin-Cheng Jhang, Fu-En Wang, Xin Yang +4

Visual grounding aims to associate free-form textual queries with specific regions in an image. While recent Multimodal Large Language Models (MLLMs) have demonstrated promising ca…

cs.CV2026

Revisiting Model Stitching In the Foundation Model Era

Zheda Mai, Ke Zhang, Fu-En Wang +6

Model stitching, connecting early layers of one model (source) to later layers of another (target) via a light stitch layer, has served as a probe of representational compatibility…

cs.CV2026

Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models

Yurou Yang, Muyuan Lin, Roberto Martin-Martin +4

Recent work explores new opportunities at the intersection of vision-language-action models (VLAs) and geometric foundation models (GFMs) for 3D reconstruction, such as VGGT. While…

cs.RO2025

Explicit Memory through Online 3D Gaussian Splatting Improves Class-Agnostic Video Segmentation

Anthony Opipari, Aravindhan K Krishnan, Shreekant Gayaka +4

Remembering where object segments were predicted in the past is useful for improving the accuracy and consistency of class-agnostic video segmentation algorithms. Existing video se…

cs.RO2025

Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning

Houjian Yu, Zheming Zhou, Min Sun +5

Enabling robots to grasp objects specified through natural language is essential for effective human-robot interaction, yet it remains a significant challenge. Existing approaches…

cs.CV2025

OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations

Peng-Hao Hsu, Ke Zhang, Fu-En Wang +6

Open-vocabulary (OV) 3D object detection is an emerging field, yet its exploration through image-based methods remains limited compared to 3D point cloud-based methods. We introduc…