11 papers
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
Gilles Puy, Nermin Samet, Alexandre Boulch +3
Plain Transformers have become the de-facto architecture for processing text, audio, image, and video, offering a unified backbone for multimodal learning. However, state-of-the-ar…
IGLOSS: Image Generation for Lidar Open-vocabulary Semantic Segmentation
Nermin Samet, Gilles Puy, Renaud Marlet
This paper presents a new method for the zero-shot open-vocabulary semantic segmentation (OVSS) of 3D automotive lidar data. To circumvent the recognized image-text modality gap th…
LOSC: LiDAR Open-voc Segmentation Consolidator
Nermin Samet, Gilles Puy, Renaud Marlet
We study the use of image-based Vision-Language Models (VLMs) for open-vocabulary segmentation of lidar scans in driving settings. Classically, image semantics can be back-projecte…
R3DPA: Leveraging 3D Representation Alignment and RGB Pretrained Priors for LiDAR Scene Generation
Nicolas Sereyjol-Garros, Ellington Kirby, Victor Besnier +1
LiDAR scene synthesis is an emerging solution to scarcity in 3D data for robotic tasks such as autonomous driving. Recent approaches employ diffusion or flow matching models to gen…
Test-Time Conditioning with Representation-Aligned Visual Features
Nicolas Sereyjol-Garros, Ellington Kirby, Victor Letzelter +2
While representation alignment with self-supervised models has been shown to improve diffusion model training, its potential for enhancing inference-time conditioning remains large…
Driving on Registers
Ellington Kirby, Alexandre Boulch, Yihong Xu +11
We present DrivoR, a simple and efficient transformer-based architecture for end-to-end autonomous driving. Our approach builds on pretrained Vision Transformers (ViTs) and introdu…