6 papers
VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining
Marwane Hariat, David Filliat, Antoine Manzanera
Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signal…
Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models
Marwane Hariat, Gianni Franchi, David Filliat +1
We present Casper3D, a lightweight probabilistic framework for converting noisy multi-view 2D foundation-model embeddings into a latent 3D semantic representation. We model view-le…
SS3D: End2End Self-Supervised 3D from Web Videos
Marwane Hariat, Gianni Franchi, David Filliat +1
We present SS3D, a web-scale SfM-based self-supervision pretraining pipeline for feed-forward 3D estimation from monocular video. Our model jointly predicts depth, ego-motion, and…
Improved monocular depth prediction using distance transform over pre-semantic contours with self-supervised neural networks
Marwane Hariat, Antoine Manzanera, David Filliat
Monocular depth estimation (MDE) with self-supervised training approaches struggles in low-texture areas, where photometric losses may lead to ambiguous depth predictions. To addre…
Double Descent Meets Out-of-Distribution Detection: Theoretical Insights and Empirical Analysis on the role of model complexity
Mouïn Ben Ammar, David Brellmann, Arturo Mendoza +2
Out-of-distribution (OOD) detection is essential for ensuring the reliability and safety of machine learning systems. In recent years, it has received increasing attention, particu…
Foundation Models and Transformers for Anomaly Detection: A Survey
Mouïn Ben Ammar, Arturo Mendoza, Nacim Belkhir +2
In line with the development of deep learning, this survey examines the transformative role of Transformers and foundation models in advancing visual anomaly detection (VAD). We ex…