6 papers
3D-Aware VLMs with Implicit and Explicit Geometries
Wenhao Li, Xueying Jiang, Quanhao Qian +4
Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial unde…
From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model
Wenhao Li, Xueying Jiang, Quanhao Qian +4
Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on actual scenarios. Existing vie…
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs
Xueying Jiang, Wenhao Li, Quanhao Qian +4
3D localization in Multimodal Large Language Models (MLLMs), including 3D object detection and 3D visual grounding, is fundamentally limited by camera intrinsic ambiguity: the same…
STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
Wenhao Li, Xueying Jiang, Gongjie Zhang +3
4D point cloud videos capture rich spatial and temporal dynamics of scenes which possess unique values in various 4D understanding tasks. However, most existing methods work in the…
Exploring 3D Reasoning-Driven Planning: From Implicit Human Intentions to Route-Aware Activity Planning
Xueying Jiang, Wenhao Li, Xiaoqin Zhang +2
3D task planning has attracted increasing attention in human-robot interaction and embodied AI thanks to the recent advances in multimodal learning. However, most existing studies…
Multimodal 3D Reasoning Segmentation with Complex Scenes
Xueying Jiang, Lewei Lu, Ling Shao +1
The recent development in multimodal learning has greatly advanced the research in 3D scene understanding in various real-world tasks such as embodied AI. However, most existing st…