collaborators

6 papers

cs.CV2026

MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos

Ziren Gong, Xiaohan Li, Fabio Tosi +4

This paper presents MAGiSt3R, a multi-agent 3D reconstruction framework performing reconstruction and camera tracking for monocular RGB videos at almost 10 FPS. MAGiSt3R relies on…

cs.RO2025

Stereo 3D Gaussian Splatting SLAM for Outdoor Urban Scenes

Xiaohan Li, Ziren Gong, Fabio Tosi +4

3D Gaussian Splatting (3DGS) has recently gained popularity in SLAM applications due to its fast rendering and high-fidelity representation. However, existing 3DGS-SLAM systems hav…

cs.CV2025

Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos

Ziren Gong, Xiaohan Li, Fabio Tosi +4

We present Ov3R, a novel framework for open-vocabulary semantic 3D reconstruction from RGB video streams, designed to advance Spatial AI. The system features two key components: CL…

cs.CV2025

DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations

Ziren Gong, Xiaohan Li, Fabio Tosi +4

This paper presents DINO-SLAM, a DINO-informed design strategy to enhance implicit (Neural Radiance Field -- NeRF) and explicit representations (Gaussian Splatting -- GS) in SLAM s…

cs.CL2025

PankRAG: Enhancing Graph Retrieval via Globally Aware Query Resolution and Dependency-Aware Reranking Mechanism

Ningyuan Li, Junrui Liu, Yi Shan +3

Recent graph-based RAG approaches leverage knowledge graphs by extracting entities from a query to fetch their associated relationships and metadata. However, relying solely on ent…

cs.CV2025

HS-SLAM: Hybrid Representation with Structural Supervision for Improved Dense SLAM

Ziren Gong, Fabio Tosi, Youmin Zhang +2

NeRF-based SLAM has recently achieved promising results in tracking and reconstruction. However, existing methods face challenges in providing sufficient scene representation, capt…