6 papers
4D-WAM: 4D Consistent World Modeling for Autonomous Driving
Jiacheng Fu, Yibo Yuan, Meng Tian +8
Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and trajectory planning. Howeve…
SUV: Future Scene Understanding as Video Generation for End-to-End Driving
Yibo Yuan, Jiacheng Fu, Jiangtong Zhu +8
End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalabilit…
AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots
Likui Zhang, Tao Tang, Zhihao Zhan +9
Recent advances in Visual-Language-Action (VLA) models have shown promising potential for robotic manipulation tasks. However, real-world robotic tasks often involve long-horizon,…
Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
Jianhua Han, Meng Tian, Jiangtong Zhu +16
Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex in…
Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning
Yue Li, Meng Tian, Dechang Zhu +4
Large vision-language models (VLMs) for autonomous driving (AD) are evolving beyond perception and cognition tasks toward motion planning. However, we identify two critical challen…
Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving
Yue Li, Meng Tian, Zhenyu Lin +7
Existing benchmarks for Vision-Language Model (VLM) on autonomous driving (AD) primarily assess interpretability through open-form visual question answering (QA) within coarse-grai…