collaborators

5 papers

cs.CV2026

OccAny: Generalized Unconstrained Urban 3D Occupancy

Anh-Quan Cao, Tuan-Hung Vu

Relying on in-domain annotations and precise sensor-rig priors, existing 3D occupancy prediction methods are limited in both scalability and out-of-domain generalization. While rec…

cs.CV2026

Vanilla ViT for Automotive Point Cloud Semantic Segmentation

Gilles Puy, Nermin Samet, Alexandre Boulch +3

Plain Transformers have become the de-facto architecture for processing text, audio, image, and video, offering a unified backbone for multimodal learning. However, state-of-the-ar…

cs.CV2026

PPT: Pretraining with Pseudo-Labeled Trajectories for Motion Forecasting

Yihong Xu, Yuan Yin, Éloi Zablocki +3

Accurately predicting how agents move in dynamic scenes is essential for safe autonomous driving. State-of-the-art motion forecasting models rely on datasets with manually annotate…

cs.CV2025

Improving Multimodal Distillation for 3D Semantic Segmentation under Domain Shift

Björn Michele, Alexandre Boulch, Gilles Puy +3

Semantic segmentation networks trained under full supervision for one type of lidar fail to generalize to unseen lidars without intervention. To reduce the performance gap under do…

cs.CV2025

VaViM and VaVAM: Autonomous Driving through Video Generative Modeling

Florent Bartoccioni, Elias Ramzi, Victor Besnier +14

We explore the potential of large-scale generative video models for autonomous driving, introducing an open-source auto-regressive video model (VaViM) and its companion video-actio…