7 papers
CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images
Haijie Li, Jiaxin Zhang, Dave Zhenyu Chen +3
Multiview image-based 3D visual grounding predicts a coordinate frame to define a coordinate system and then regresses a 3D bounding box for localization. However, existing methods…
Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction
Zecheng Tang, Jiaye Fu, Qiankun Gao +5
Feed-forward 3D reconstruction models based on Vision Transformers can directly estimate scene geometry and camera poses from a small set of input images, but scaling them to video…
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
Jiaxin Zhang, Junjun Jiang, Haijie Li +3
Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging i…
InstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perception
Haijie Li, Yanmin Wu, Jiarui Meng +4
3D scene understanding has become an essential area of research with applications in autonomous driving, robotics, and augmented reality. Recently, 3D Gaussian Splatting (3DGS) has…
Mirror-3DGS: Incorporating Mirror Reflections into 3D Gaussian Splatting
Jiarui Meng, Haijie Li, Yanmin Wu +4
3D Gaussian Splatting (3DGS) has significantly advanced 3D scene reconstruction and novel view synthesis. However, like Neural Radiance Fields (NeRF), 3DGS struggles with accuratel…
OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding
Yanmin Wu, Jiarui Meng, Haijie Li +8
This paper introduces OpenGaussian, a method based on 3D Gaussian Splatting (3DGS) capable of 3D point-level open vocabulary understanding. Our primary motivation stems from observ…