6 papers
MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos
Ziren Gong, Xiaohan Li, Fabio Tosi +4
This paper presents MAGiSt3R, a multi-agent 3D reconstruction framework performing reconstruction and camera tracking for monocular RGB videos at almost 10 FPS. MAGiSt3R relies on…
Stereo 3D Gaussian Splatting SLAM for Outdoor Urban Scenes
Xiaohan Li, Ziren Gong, Fabio Tosi +4
3D Gaussian Splatting (3DGS) has recently gained popularity in SLAM applications due to its fast rendering and high-fidelity representation. However, existing 3DGS-SLAM systems hav…
Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
Ziren Gong, Xiaohan Li, Fabio Tosi +4
We present Ov3R, a novel framework for open-vocabulary semantic 3D reconstruction from RGB video streams, designed to advance Spatial AI. The system features two key components: CL…
DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations
Ziren Gong, Xiaohan Li, Fabio Tosi +4
This paper presents DINO-SLAM, a DINO-informed design strategy to enhance implicit (Neural Radiance Field -- NeRF) and explicit representations (Gaussian Splatting -- GS) in SLAM s…
PankRAG: Enhancing Graph Retrieval via Globally Aware Query Resolution and Dependency-Aware Reranking Mechanism
Ningyuan Li, Junrui Liu, Yi Shan +3
Recent graph-based RAG approaches leverage knowledge graphs by extracting entities from a query to fetch their associated relationships and metadata. However, relying solely on ent…
HS-SLAM: Hybrid Representation with Structural Supervision for Improved Dense SLAM
Ziren Gong, Fabio Tosi, Youmin Zhang +2
NeRF-based SLAM has recently achieved promising results in tracking and reconstruction. However, existing methods face challenges in providing sufficient scene representation, capt…