4 papers
Enhancing Part-Level Point Grounding for Any Open-Source MLLMs
Jin-Cheng Jhang, Fu-En Wang, Xin Yang +4
Visual grounding aims to associate free-form textual queries with specific regions in an image. While recent Multimodal Large Language Models (MLLMs) have demonstrated promising ca…
Direct Action-Head Injection of A Grounded 3D Point Unlocks Spatial and Task Generalization
Shiang-Feng Tsai, Jin-Cheng Jhang, Yen-Ling Tai +4
Vision-Language-Action (VLA) models leverage large-scale vision-language pretraining for flexible robot manipulation, yet at test time they remain brittle along two axes: spatial g…
uLayout: Unified Room Layout Estimation for Perspective and Panoramic Images
Jonathan Lee, Bolivar Solarte, Chin-Hsuan Wu +4
We present uLayout, a unified model for estimating room layout geometries from both perspective and panoramic images, whereas traditional solutions require different model designs…
V-MIND: Building Versatile Monocular Indoor 3D Detector with Diverse 2D Annotations
Jin-Cheng Jhang, Tao Tu, Fu-En Wang +3
The field of indoor monocular 3D object detection is gaining significant attention, fueled by the increasing demand in VR/AR and robotic applications. However, its advancement is i…