1 citations · 1 across the 12 of their papers we have counts for
4 papers · 2 filters
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training
Yiyang Du, Zhanqiu Guo, Xin Ye +2
Vision-Language-Action Models (VLAs) inherit their visual and linguistic capabilities from Vision-Language Models (VLMs), yet most VLAs are built from off-the-shelf VLMs that are n…
ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving
Zihao Sheng, Xin Ye, Jingru Luo +2
End-to-end autonomous driving models based on Vision-Language-Action (VLA) architectures have shown promising results by learning driving policies through behavior cloning on exper…
Reconstruction Matters: Learning Geometry-Aligned BEV Representation through 3D Gaussian Splatting
Yiren Lu, Xin Ye, Burhaneddin Yaman +4
Bird's-Eye-View (BEV) perception serves as a cornerstone for autonomous driving, offering a unified spatial representation that fuses surrounding-view images to enable reasoning fo…
UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving
Zhexiao Xiong, Xin Ye, Burhan Yaman +5
World models have become central to autonomous driving, where accurate scene understanding and future prediction are crucial for safe control. Recent work has explored using vision…