8 papers
Exploring Easy Boosts for Lidar Semantic Scene Completion
Tetiana Martyniuk, Jonathan Seele, Alexandre Boulch +3
This paper investigates "free lunch" strategies to boost the performance of lidar semantic scene completion (SSC) without requiring complex architectural redesigns. We first demons…
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
Gilles Puy, Nermin Samet, Alexandre Boulch +3
Plain Transformers have become the de-facto architecture for processing text, audio, image, and video, offering a unified backbone for multimodal learning. However, state-of-the-ar…
IGLOSS: Image Generation for Lidar Open-vocabulary Semantic Segmentation
Nermin Samet, Gilles Puy, Renaud Marlet
This paper presents a new method for the zero-shot open-vocabulary semantic segmentation (OVSS) of 3D automotive lidar data. To circumvent the recognized image-text modality gap th…
3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
Ryousuke Yamada, Kohsuke Ide, Yoshihiro Fukuhara +4
Despite recent progress in 3D self-supervised learning, collecting large-scale 3D scene scans remains expensive and labor-intensive. In this work, we investigate whether 3D represe…
Driving on Registers
Ellington Kirby, Alexandre Boulch, Yihong Xu +11
We present DrivoR, a simple and efficient transformer-based architecture for end-to-end autonomous driving. Our approach builds on pretrained Vision Transformers (ViTs) and introdu…
Improving Multimodal Distillation for 3D Semantic Segmentation under Domain Shift
Björn Michele, Alexandre Boulch, Gilles Puy +3
Semantic segmentation networks trained under full supervision for one type of lidar fail to generalize to unseen lidars without intervention. To reduce the performance gap under do…