Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models
Haibo Wang, Lifu Huang
Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations that fail to maintain geometric and spatia…
cs.CV2026
Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
Haibo Wang, Zihao Lin, Zhiyang Xu +1
3D Visual Grounding (3D-VG) aims to localize objects in 3D scenes via natural language descriptions. While recent advancements leveraging Vision-Language Models (VLMs) have explore…