4 papers · 1 filter
Towards Zero-Shot Transfer Across Embodiments For Driving VLAs
Caio Azevedo, Stefano Sabatini, Sascha Hornauer +1
Vision-Language-Action models (VLAs) have shown strong potential in autonomous driving by leveraging multimodal pretraining for instruction following, visual reasoning, and scene-l…
LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding
Fusang Wang, Nathan Piasco, Moussab Bennehar +3
Recent advancements in open-vocabulary 3D scene understanding heavily rely on 3D Gaussian Splatting (3DGS) to register vision-language features into 3D space. However, we identify…
LiDAS: Lighting-driven Dynamic Active Sensing for Nighttime Perception
Simon de Moreau, Andrei Bursuc, Hafid El-Idrissi +1
Nighttime environments pose significant challenges for camera-based perception, as existing methods passively rely on the scene lighting. We introduce Lighting-driven Dynamic Activ…
AVS-Net: Audio-Visual Scale Net for Self-supervised Monocular Metric Depth Estimation
Xiaohu Liu, Sascha Hornauer, Fabien Moutarde +1
Metric depth prediction from monocular videos suffers from bad generalization between datasets and requires supervised depth data for scale-correct training. Self-supervised traini…