6 papers
CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images
Haijie Li, Jiaxin Zhang, Dave Zhenyu Chen +3
Multiview image-based 3D visual grounding predicts a coordinate frame to define a coordinate system and then regresses a 3D bounding box for localization. However, existing methods…
Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction
Zecheng Tang, Jiaye Fu, Qiankun Gao +5
Feed-forward 3D reconstruction models based on Vision Transformers can directly estimate scene geometry and camera poses from a small set of input images, but scaling them to video…
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
Jiaxin Zhang, Junjun Jiang, Haijie Li +3
Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging i…
InstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perception
Haijie Li, Yanmin Wu, Jiarui Meng +4
3D scene understanding has become an essential area of research with applications in autonomous driving, robotics, and augmented reality. Recently, 3D Gaussian Splatting (3DGS) has…
OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding
Yanmin Wu, Jiarui Meng, Haijie Li +8
This paper introduces OpenGaussian, a method based on 3D Gaussian Splatting (3DGS) capable of 3D point-level open vocabulary understanding. Our primary motivation stems from observ…
Hybrid Fourier Score Distillation for Efficient One Image to 3D Object Generation
Shuzhou Yang, Yu Wang, Haijie Li +4
Single image-to-3D generation is pivotal for crafting controllable 3D assets. Given its under-constrained nature, we attempt to leverage 3D geometric priors from a novel view diffu…