7 papers
DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception
Tim Broedermannn, Christos Sakaridis, Luigi Piccinelli +2
Robust semantic perception for autonomous vehicles relies on effectively combining multiple sensors with complementary strengths and weaknesses. State-of-the-art sensor fusion appr…
ACDC: The Adverse Conditions Dataset with Correspondences for Robust Semantic Driving Scene Perception
Christos Sakaridis, Haoran Wang, Ke Li +6
Level-5 driving automation requires a robust visual perception system that can parse input images under any condition. However, existing driving datasets for dense semantic percept…
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
Hao Tang, Ling Shao, Zhenyu Zhang +2
We propose a novel spatial-temporal graph Mamba (STG-Mamba) for the music-guided dance video synthesis task, i.e., to translate the input music to a dance video. STG-Mamba consists…
GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond
Anna-Maria Halacheva, Jan-Nico Zaech, Xi Wang +2
As multimodal language models advance, their application to 3D scene understanding is a fast-growing frontier, driving the development of 3D Vision-Language Models (VLMs). Current…
CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes
Tim Broedermann, Christos Sakaridis, Yuqian Fu +1
Leveraging multiple sensors is crucial for robust semantic perception in autonomous driving, as each sensor type has complementary strengths and weaknesses. However, existing senso…
Condition-Invariant Semantic Segmentation
Christos Sakaridis, David Bruggemann, Fisher Yu +1
Adaptation of semantic segmentation networks to different visual conditions is vital for robust perception in autonomous cars and robots. However, previous work has shown that most…