4 papers · 2 filters
GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
Hao Li, Han Fang, Zixin Pan +8
Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge. Existing me…
Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
Tianyi Gao, Han Fang, Tianyi Ding +9
Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing…
DrivingDepth: Sparse-Prompted Pixel-wise Scale Correction for Driving Depth Estimation
Chi Huang, Wenhao Zhang, Hang Yin +5
Dense depth estimation for autonomous driving faces a geometry-scale conflict: depth foundation models deliver pixel-aligned dense visual geometry without reliable metric scale, wh…
Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting
Xiaobiao Du, YuAn Wang, Hao Li +3
Recent advances in 3D Gaussian Splatting have demonstrated unprecedented success in novel view synthesis. However, the substantial inference and storage overhead driven by high-ord…