5 papers
Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction
Junhong Lin, Jinlong Wang, Xianda Guo +6
Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing methods rely on complex dec…
Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning
Qianlong Yang, Bowen Ye, Xianda Guo +4
Despite the progress of multimodal large language models (MLLMs), they continue to exhibit deficiencies in visual perception. Following visual instruction tuning, internal MLLM rep…
VGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy Prediction
Junhong Lin, Xianda Guo, Kangli Wang +4
Vision-only occupancy prediction requires recovering a semantic 3D occupancy field from calibrated surround-view images, where each view provides observations with ambiguous depth…
Bridging 3D Gaussians and Semantic Occupancy for Comprehensive Open-Vocabulary Scene Understanding from Unposed Images
Hu Zhu, Bohan Li, Xianda Guo +5
Comprehensive 3D scene understanding from sparse, unposed images requires a model to recover renderable geometry, open-vocabulary semantics, and free/occupied 3D space without rely…
Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading
Hong Li, Minqi Meng, Yanjun Liang +10
Reconstructing high-fidelity, relightable 3D avatars from a single in-the-wild image is a challenging ill-posed problem, primarily hindered by the scarcity of high-quality PBR data…