5 papers
OccAny: Generalized Unconstrained Urban 3D Occupancy
Anh-Quan Cao, Tuan-Hung Vu
Relying on in-domain annotations and precise sensor-rig priors, existing 3D occupancy prediction methods are limited in both scalability and out-of-domain generalization. While rec…
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
Gilles Puy, Nermin Samet, Alexandre Boulch +3
Plain Transformers have become the de-facto architecture for processing text, audio, image, and video, offering a unified backbone for multimodal learning. However, state-of-the-ar…
PPT: Pretraining with Pseudo-Labeled Trajectories for Motion Forecasting
Yihong Xu, Yuan Yin, Ãloi Zablocki +3
Accurately predicting how agents move in dynamic scenes is essential for safe autonomous driving. State-of-the-art motion forecasting models rely on datasets with manually annotate…
Improving Multimodal Distillation for 3D Semantic Segmentation under Domain Shift
Björn Michele, Alexandre Boulch, Gilles Puy +3
Semantic segmentation networks trained under full supervision for one type of lidar fail to generalize to unseen lidars without intervention. To reduce the performance gap under do…
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
Florent Bartoccioni, Elias Ramzi, Victor Besnier +14
We explore the potential of large-scale generative video models for autonomous driving, introducing an open-source auto-regressive video model (VaViM) and its companion video-actio…