6 papers
VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining
Marwane Hariat, David Filliat, Antoine Manzanera
Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signal…
Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models
Marwane Hariat, Gianni Franchi, David Filliat +1
We present Casper3D, a lightweight probabilistic framework for converting noisy multi-view 2D foundation-model embeddings into a latent 3D semantic representation. We model view-le…
SS3D: End2End Self-Supervised 3D from Web Videos
Marwane Hariat, Gianni Franchi, David Filliat +1
We present SS3D, a web-scale SfM-based self-supervision pretraining pipeline for feed-forward 3D estimation from monocular video. Our model jointly predicts depth, ego-motion, and…
Improved monocular depth prediction using distance transform over pre-semantic contours with self-supervised neural networks
Marwane Hariat, Antoine Manzanera, David Filliat
Monocular depth estimation (MDE) with self-supervised training approaches struggles in low-texture areas, where photometric losses may lead to ambiguous depth predictions. To addre…
Rebalancing gradient to improve self-supervised co-training of depth, odometry and optical flow predictions
Marwane Hariat, Antoine Manzanera, David Filliat
We present CoopNet, an approach that improves the cooperation of co-trained networks by dynamically adapting the apportionment of gradient, to ensure equitable learning progress. I…
Hierarchical Light Transformer Ensembles for Multimodal Trajectory Forecasting
Adrien Lafage, Mathieu Barbier, Gianni Franchi +1
Accurate trajectory forecasting is crucial for the performance of various systems, such as advanced driver-assistance systems and self-driving vehicles. These forecasts allow us to…