activity
20242026
collaborators

6 papers

cs.RO2026

A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world

Steeven Janny, Leonid Antsfeld, Christian Wolf

Trained policies for real-world robotics rely on computer vision components, typically in the form of pre-trained visual encoders. These encoders are an essential component and it…

cs.CV2026

S-MUSt3R: Sliding Multi-view 3D Reconstruction

Leonid Antsfeld, Boris Chidlovskii, Yohann Cabon +2

The recent paradigm shift in 3D vision led to the rise of foundation models with remarkable capabilities in 3D perception from uncalibrated images. However, extending these models…

cs.CV2025

PanSt3R: Multi-view Consistent Panoptic Segmentation

Lojze Zust, Yohann Cabon, Juliette Marrie +4

Panoptic segmentation of 3D scenes, involving the segmentation and classification of object instances in a dense 3D reconstruction of a scene, is a challenging problem, especially…

cs.RO2025

Reasoning in visual navigation of end-to-end trained agents: a dynamical systems approach

Steeven Janny, Hervé Poirier, Leonid Antsfeld +6

Progress in Embodied AI has made it possible for end-to-end-trained agents to navigate in photo-realistic environments with high-level reasoning and zero-shot or language-condition…

cs.CV2025

MUSt3R: Multi-view Network for Stereo 3D Reconstruction

Yohann Cabon, Lucas Stoffl, Leonid Antsfeld +4

DUSt3R introduced a novel paradigm in geometric computer vision by proposing a model that can provide dense and unconstrained Stereo 3D Reconstruction of arbitrary image collection…

cs.CV2024

3D-Consistent Image Inpainting with Diffusion Models

Leonid Antsfeld, Boris Chidlovskii

We address the problem of 3D inconsistency of image inpainting based on diffusion models. We propose a generative model using image pairs that belong to the same scene. To achieve…