7 papers · 1 filter
How Can Driving World Models Do Counterfactual Prediction?
Jiaru Zhang, Can Cui, Yi Xu +3
Driving world models are often interpreted as counterfactual simulators for observed driving episodes: given a factual driving log, they are asked what would have happened under an…
UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving
Zhexiao Xiong, Xin Ye, Burhan Yaman +5
World models have become central to autonomous driving, where accurate scene understanding and future prediction are crucial for safe control. Recent work has explored using vision…
ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving
Zihao Sheng, Xin Ye, Jingru Luo +2
End-to-end autonomous driving models based on Vision-Language-Action (VLA) architectures have shown promising results by learning driving policies through behavior cloning on exper…
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training
Yiyang Du, Zhanqiu Guo, Xin Ye +2
Vision-Language-Action Models (VLAs) inherit their visual and linguistic capabilities from Vision-Language Models (VLMs), yet most VLAs are built from off-the-shelf VLMs that are n…
ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving
Yunsheng Ma, Burhaneddin Yaman, Xin Ye +5
Recent advances have explored integrating large language models (LLMs) into end-to-end autonomous driving systems to enhance generalization and interpretability. However, most exis…
BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance
Xin Ye, Burhaneddin Yaman, Sheng Cheng +3
Bird's-eye-view (BEV) representations play a crucial role in autonomous driving tasks. Despite recent advancements in BEV generation, inherent noise, stemming from sensor limitatio…