activity
20242026
collaborators

6 papers

cs.CV2026

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs

Jin-Cheng Jhang, Fu-En Wang, Xin Yang +4

Visual grounding aims to associate free-form textual queries with specific regions in an image. While recent Multimodal Large Language Models (MLLMs) have demonstrated promising ca…

cs.CV2026

Revisiting Model Stitching In the Foundation Model Era

Zheda Mai, Ke Zhang, Fu-En Wang +6

Model stitching, connecting early layers of one model (source) to later layers of another (target) via a light stitch layer, has served as a probe of representational compatibility…

cs.CV2025

OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations

Peng-Hao Hsu, Ke Zhang, Fu-En Wang +6

Open-vocabulary (OV) 3D object detection is an emerging field, yet its exploration through image-based methods remains limited compared to 3D point cloud-based methods. We introduc…

cs.CV2025

UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References

Ming-Feng Li, Xin Yang, Fu-En Wang +5

6D object pose estimation has shown strong generalizability to novel objects. However, existing methods often require either a complete, well-reconstructed 3D model or numerous ref…

cs.CV2025

uLayout: Unified Room Layout Estimation for Perspective and Panoramic Images

Jonathan Lee, Bolivar Solarte, Chin-Hsuan Wu +4

We present uLayout, a unified model for estimating room layout geometries from both perspective and panoramic images, whereas traditional solutions require different model designs…

cs.CV2024

V-MIND: Building Versatile Monocular Indoor 3D Detector with Diverse 2D Annotations

Jin-Cheng Jhang, Tao Tu, Fu-En Wang +3

The field of indoor monocular 3D object detection is gaining significant attention, fueled by the increasing demand in VR/AR and robotic applications. However, its advancement is i…