3 papers
cs.CV2025
Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
Meng Cao, Haokun Lin, Haoyuan Li +6
Spatial reasoning, the ability to understand and interpret the 3D structure of the world, is a critical yet underdeveloped capability in Multimodal Large Language Models (MLLMs). C…
cs.CV2025
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
Haoyuan Li, Yanpeng Zhou, Yufei Gao +7
Remarkable progress in 2D Vision-Language Models (VLMs) has spurred interest in extending them to 3D settings for tasks like 3D Question Answering, Dense Captioning, and Visual Gro…
cs.CV2025
UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting
Haoyuan Li, Yanpeng Zhou, Tao Tang +5
Recent advancements in multi-modal 3D pre-training methods have shown promising efficacy in learning joint representations of text, images, and point clouds. However, adopting poin…