From the 1 of 8 linked papers with an AI index.
4 papers · 1 filter
MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight
Zehua Fan, Junjie He, Wenxuan Song +14
World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands sim…
FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models
Wei Li, Peijin Jia, Yuan Ma +9
FoMoVLA enhances vision-language-action models by jointly predicting future visual features and tracking sparse 2D points, providing both goal states and motion paths to improve co…
Unifying Language-Action Understanding and Generation for Autonomous Driving
Xinyang Wang, Qian Liu, Wenjie Ding +7
Vision-Language-Action (VLA) models are emerging as a promising paradigm for end-to-end autonomous driving, valued for their potential to leverage world knowledge and reason about…
DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
Xiaoyu Tian, Junru Gu, Bailin Li +7
A primary hurdle of autonomous driving in urban environments is understanding complex and long-tail scenarios, such as challenging road conditions and delicate human behaviors. We…