From the 2 of 19 linked papers with an AI index.
19 papers
MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos
Ziren Gong, Xiaohan Li, Fabio Tosi +4
The paper introduces MAGiSt3R, a multi-agent framework that reconstructs 3D scenes and tracks camera pose from monocular RGB videos in near real-time using feed-forward models and…
DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations
Ziren Gong, Xiaohan Li, Fabio Tosi +4
The paper introduces DINO-SLAM, a system that combines DINO semantic features with a geometry encoder to improve both neural implicit (NeRF) and explicit (Gaussian Splatting) SLAM…
ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device
Fabio Tosi, Luca Bartolomei, Matteo Poggi +1
Monocular depth estimation has seen remarkable progress through foundation models achieving robust zero-shot generalization, yet their computational demands place them far beyond t…
FlowIt: Global Matching via Hierarchical Transformers and Optimal Transport for Optical Flow
Sadra Safadoust, Fabio Tosi, Matteo Poggi +1
We present FlowIt, a novel architecture for optical flow estimation that combines global matching with confidence and occlusion-guided refinement. At its core, FlowIt leverages a h…
StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space
Tjark Behrens, Anton Obukhov, Bingxin Ke +3
We introduce StereoSpace, a diffusion-based framework for monocular-to-stereo synthesis that models geometry purely through viewpoint conditioning, without explicit depth or warpin…
Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo
Ninghui Xu, Fabio Tosi, Lihui Wang +5
Conventional frame-based cameras capture rich contextual information but suffer from limited temporal resolution and motion blur in dynamic scenes. Event cameras offer an alternati…