4 papers
MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos
Ziren Gong, Xiaohan Li, Fabio Tosi +4
This paper presents MAGiSt3R, a multi-agent 3D reconstruction framework performing reconstruction and camera tracking for monocular RGB videos at almost 10 FPS. MAGiSt3R relies on…
Stereo 3D Gaussian Splatting SLAM for Outdoor Urban Scenes
Xiaohan Li, Ziren Gong, Fabio Tosi +4
3D Gaussian Splatting (3DGS) has recently gained popularity in SLAM applications due to its fast rendering and high-fidelity representation. However, existing 3DGS-SLAM systems hav…
Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
Ziren Gong, Xiaohan Li, Fabio Tosi +4
We present Ov3R, a novel framework for open-vocabulary semantic 3D reconstruction from RGB video streams, designed to advance Spatial AI. The system features two key components: CL…
DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations
Ziren Gong, Xiaohan Li, Fabio Tosi +4
This paper presents DINO-SLAM, a DINO-informed design strategy to enhance implicit (Neural Radiance Field -- NeRF) and explicit representations (Gaussian Splatting -- GS) in SLAM s…