4 papers · 1 filter
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training
Yiyang Du, Zhanqiu Guo, Xin Ye +2
Vision-Language-Action Models (VLAs) inherit their visual and linguistic capabilities from Vision-Language Models (VLMs), yet most VLAs are built from off-the-shelf VLMs that are n…
ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving
Yunsheng Ma, Burhaneddin Yaman, Xin Ye +5
Recent advances have explored integrating large language models (LLMs) into end-to-end autonomous driving systems to enhance generalization and interpretability. However, most exis…
BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance
Xin Ye, Burhaneddin Yaman, Sheng Cheng +3
Bird's-eye-view (BEV) representations play a crucial role in autonomous driving tasks. Despite recent advancements in BEV generation, inherent noise, stemming from sensor limitatio…
MTA: Multimodal Task Alignment for BEV Perception and Captioning
Yunsheng Ma, Burhaneddin Yaman, Xin Ye +5
Bird's eye view (BEV)-based 3D perception plays a crucial role in autonomous driving applications. The rise of large language models has spurred interest in BEV-based captioning to…