From the 1 of 8 linked papers with an AI index.
8 papers
CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images
Haijie Li, Jiaxin Zhang, Dave Zhenyu Chen +3
Multiview image-based 3D visual grounding predicts a coordinate frame to define a coordinate system and then regresses a 3D bounding box for localization. However, existing methods…
AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling
Yingji Zhong, Dave Zhenyu Chen, Fuzhao Ou +4
The paper introduces AsySplat, an asymmetric architecture that separates geometry and appearance modeling in 3D Gaussian splatting to reduce redundant computation and achieve fast,…
Reliev3R: Relieving Feed-forward Reconstruction from Multi-View Geometric Annotations
Youyu Chen, Junjun Jiang, Yueru Luo +4
With recent advances, Feed-forward Reconstruction Models (FFRMs) have demonstrated great potential in reconstruction quality and adaptiveness to multiple downstream tasks. However,…
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
Jiaxin Zhang, Junjun Jiang, Haijie Li +3
Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging i…
FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust Fusion
Pihai Sun, Junjun Jiang, Yuanqi Yao +4
Image-event joint depth estimation methods leverage complementary modalities for robust perception, yet face challenges in generalizability stemming from two factors: 1) limited an…
Quantifying and Alleviating Co-Adaptation in Sparse-View 3D Gaussian Splatting
Kangjie Chen, Yingji Zhong, Zhihao Li +4
3D Gaussian Splatting (3DGS) has demonstrated impressive performance in novel view synthesis under dense-view settings. However, in sparse-view scenarios, despite the realistic ren…