collaborators

6 papers

cs.CV2026

3D-Aware VLMs with Implicit and Explicit Geometries

Wenhao Li, Xueying Jiang, Quanhao Qian +4

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial unde…

cs.CV2026

From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model

Wenhao Li, Xueying Jiang, Quanhao Qian +4

Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting depending on actual scenarios. Existing vie…

cs.CV2026

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs

Xueying Jiang, Wenhao Li, Quanhao Qian +4

3D localization in Multimodal Large Language Models (MLLMs), including 3D object detection and 3D visual grounding, is fundamentally limited by camera intrinsic ambiguity: the same…

cs.CV2026

STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding

Wenhao Li, Xueying Jiang, Gongjie Zhang +3

4D point cloud videos capture rich spatial and temporal dynamics of scenes which possess unique values in various 4D understanding tasks. However, most existing methods work in the…

cs.CV2025

Exploring 3D Reasoning-Driven Planning: From Implicit Human Intentions to Route-Aware Activity Planning

Xueying Jiang, Wenhao Li, Xiaoqin Zhang +2

3D task planning has attracted increasing attention in human-robot interaction and embodied AI thanks to the recent advances in multimodal learning. However, most existing studies…

cs.CV2025

Multimodal 3D Reasoning Segmentation with Complex Scenes

Xueying Jiang, Lewei Lu, Ling Shao +1

The recent development in multimodal learning has greatly advanced the research in 3D scene understanding in various real-world tasks such as embodied AI. However, most existing st…