5 papers
One Agent to Guide Them All: Empowering MLLMs for Vision-and-Language Navigation via Explicit World Representation
Zerui Li, Hongpei Zheng, Fangguo Zhao +5
A navigable agent needs to understand both high-level semantic instructions and precise spatial perceptions. Building navigation agents centered on Multimodal Large Language Models…
SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding
Hongpei Zheng, Shijie Li, Yanran Li +1
Spatial reasoning in large-scale 3D environments remains challenging for current vision-language models, which are typically constrained to room-scale scenarios. We introduce H…
Reg3D: Reconstructive Geometry Instruction Tuning for 3D Scene Understanding
Hongpei Zheng, Lintao Xiang, Qijun Yang +2
The rapid development of Large Multimodal Models (LMMs) has led to remarkable progress in 2D visual understanding; however, extending these capabilities to 3D scene understanding r…
PointGS: Point Attention-Aware Sparse View Synthesis with Gaussian Splatting
Lintao Xiang, Hongpei Zheng, Yating Huang +2
3D Gaussian splatting (3DGS) is an innovative rendering technique that surpasses the neural radiance field (NeRF) in both rendering speed and visual quality by leveraging an explic…
Geometric Prior-Guided Neural Implicit Surface Reconstruction in the Wild
Lintao Xiang, Hongpei Zheng, Bailin Deng +1
Neural implicit surface reconstruction using volume rendering techniques has recently achieved significant advancements in creating high-fidelity surfaces from multiple 2D images.…