4 papers
PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth
Bu Jin, Weize Li, Baihan Yang +9
Recent advancements in autonomous driving (AD) systems have highlighted the potential of world models in achieving robust and generalizable performance across both ordinary and cha…
MonoOcc: Digging into Monocular Semantic Occupancy Prediction
Yupeng Zheng, Xiang Li, Pengfei Li +6
Monocular Semantic Occupancy Prediction aims to infer the complete 3D geometry and semantic information of scenes from only 2D images. It has garnered significant attention, partic…
TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
Bu Jin, Yupeng Zheng, Pengfei Li +12
3D dense captioning stands as a cornerstone in achieving a comprehensive understanding of 3D scenes through natural language. It has recently witnessed remarkable achievements, par…
Large Language Models Powered Context-aware Motion Prediction in Autonomous Driving
Xiaoji Zheng, Lixiu Wu, Zhijie Yan +5
Motion prediction is among the most fundamental tasks in autonomous driving. Traditional methods of motion forecasting primarily encode vector information of maps and historical tr…