4 papers
CLIP-GS: CLIP-Informed Gaussian Splatting for View-Consistent 3D Indoor Semantic Understanding
Guibiao Liao, Jiankun Li, Zhenyu Bao +3
Exploiting 3D Gaussian Splatting (3DGS) with Contrastive Language-Image Pre-Training (CLIP) models for open-vocabulary 3D semantic understanding of indoor scenes has emerged as an…
BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents
Yumeng Zhang, Shi Gong, Kaixin Xiong +7
World models have attracted increasing attention in autonomous driving for their ability to forecast potential future scenarios. In this paper, we propose BEVWorld, a novel framewo…
Exploring the Causality of End-to-End Autonomous Driving
Jiankun Li, Hao Li, Jiangjiang Liu +6
Deep learning-based models are widely deployed in autonomous driving areas, especially the increasingly noticed end-to-end solutions. However, the black-box property of these model…
BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-based Roadside 3D Object Detection
Wenjie Wang, Yehao Lu, Guangcong Zheng +6
Vision-based roadside 3D object detection has attracted rising attention in autonomous driving domain, since it encompasses inherent advantages in reducing blind spots and expandin…