works on

From the 2 of 15 linked papers with an AI index.

collaborators

15 papers

cs.CV2026

MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos

Ziren Gong, Xiaohan Li, Fabio Tosi +4

The paper introduces MAGiSt3R, a multi-agent framework that reconstructs 3D scenes and tracks camera pose from monocular RGB videos in near real-time using feed-forward models and…

cs.CV2026

DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations

Ziren Gong, Xiaohan Li, Fabio Tosi +4

The paper introduces DINO-SLAM, a system that combines DINO semantic features with a geometry encoder to improve both neural implicit (NeRF) and explicit (Gaussian Splatting) SLAM…

cs.CV2026

ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device

Fabio Tosi, Luca Bartolomei, Matteo Poggi +1

Monocular depth estimation has seen remarkable progress through foundation models achieving robust zero-shot generalization, yet their computational demands place them far beyond t…

cs.CV2026

Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo

Ninghui Xu, Fabio Tosi, Lihui Wang +5

Conventional frame-based cameras capture rich contextual information but suffer from limited temporal resolution and motion blur in dynamic scenes. Event cameras offer an alternati…

cs.CV2026

EventHub: Data Factory for Generalizable Event-Based Stereo Networks without Active Sensors

Luca Bartolomei, Fabio Tosi, Matteo Poggi +2

We propose EventHub, a novel framework for training deep-event stereo networks without ground truth annotations from costly active sensors, relying instead on standard color images…

cs.CV2025

Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos

Ziren Gong, Xiaohan Li, Fabio Tosi +4

We present Ov3R, a novel framework for open-vocabulary semantic 3D reconstruction from RGB video streams, designed to advance Spatial AI. The system features two key components: CL…