8 papers
3D-Aware VLMs with Implicit and Explicit Geometries
Wenhao Li, Xueying Jiang, Quanhao Qian +4
Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial unde…
GeoProp: Grounding Robot State in Vision for Generalist Manipulation
Guoyang Zhao, Quanhao Qian, Gongjie Zhang +5
Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a dir…
From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model
Wenhao Li, Xueying Jiang, Quanhao Qian +4
Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on actual scenarios. Existing vie…
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs
Xueying Jiang, Wenhao Li, Quanhao Qian +4
3D localization in Multimodal Large Language Models (MLLMs), including 3D object detection and 3D visual grounding, is fundamentally limited by camera intrinsic ambiguity: the same…
On the Generalization Capacities of MLLMs for Spatial Intelligence
Gongjie Zhang, Wenhao Li, Quanhao Qian +4
Multimodal Large Language Models (MLLMs) that directly process RGB inputs for tasks like 3D localization and navigation have shown remarkable potential. However, we argue that thes…
RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
Jiuniu Wang, Gongjie Zhang, Quanhao Qian +3
Scalable Vector Graphics (SVGs) are fundamental to digital design and robot control, encoding not only visual structure but also motion paths in interactive drawings. In this work,…