activity
20242026
collaborators

7 papers

cs.CV2026

CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images

Haijie Li, Jiaxin Zhang, Dave Zhenyu Chen +3

Multiview image-based 3D visual grounding predicts a coordinate frame to define a coordinate system and then regresses a 3D bounding box for localization. However, existing methods…

cs.CV2026

Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction

Zecheng Tang, Jiaye Fu, Qiankun Gao +5

Feed-forward 3D reconstruction models based on Vision Transformers can directly estimate scene geometry and camera poses from a small set of input images, but scaling them to video…

cs.CV2026

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models

Jiaxin Zhang, Junjun Jiang, Haijie Li +3

Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging i…

cs.CV2025

InstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perception

Haijie Li, Yanmin Wu, Jiarui Meng +4

3D scene understanding has become an essential area of research with applications in autonomous driving, robotics, and augmented reality. Recently, 3D Gaussian Splatting (3DGS) has…

cs.CV2024

Mirror-3DGS: Incorporating Mirror Reflections into 3D Gaussian Splatting

Jiarui Meng, Haijie Li, Yanmin Wu +4

3D Gaussian Splatting (3DGS) has significantly advanced 3D scene reconstruction and novel view synthesis. However, like Neural Radiance Fields (NeRF), 3DGS struggles with accuratel…

cs.CV2024

OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding

Yanmin Wu, Jiarui Meng, Haijie Li +8

This paper introduces OpenGaussian, a method based on 3D Gaussian Splatting (3DGS) capable of 3D point-level open vocabulary understanding. Our primary motivation stems from observ…