14 papers
CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images
Haijie Li, Jiaxin Zhang, Dave Zhenyu Chen +3
Multiview image-based 3D visual grounding predicts a coordinate frame to define a coordinate system and then regresses a 3D bounding box for localization. However, existing methods…
Breaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation
Kaihua Tang, Ziqing Xia, Xiaoxu Zheng +4
Despite recent advances in Monocular Depth Estimation, state-of-the-art depth foundation models remain vulnerable to robustness issues. Particularly, even slight camera rolls can r…
AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling
Yingji Zhong, Dave Zhenyu Chen, Fuzhao Ou +4
The paper introduces AsySplat, an asymmetric architecture that separates geometry and appearance modeling in 3D Gaussian splatting to reduce redundant computation and achieve fast,…
OneCanvas: 3D Scene Understanding via Panoramic Reprojection
BartÅomiej Baranowski, Dave Zhenyu Chen, Matthias NieÃner
Existing approaches to 3D scene understanding in Vision-Language Models (VLMs) either rely on complex, model-specific geometry encoders or large training budgets in pursuit of spat…
AnchorSplat: Feed-Forward 3D Gaussian Splatting with 3D Geometric Priors
Xiaoxue Zhang, Xiaoxu Zheng, Yixuan Yin +5
Recent feed-forward Gaussian reconstruction models adopt a pixel-aligned formulation that maps each 2D pixel to a 3D Gaussian, entangling Gaussian representations tightly with the…
Reliev3R: Relieving Feed-forward Reconstruction from Multi-View Geometric Annotations
Youyu Chen, Junjun Jiang, Yueru Luo +4
With recent advances, Feed-forward Reconstruction Models (FFRMs) have demonstrated great potential in reconstruction quality and adaptiveness to multiple downstream tasks. However,…