collaborators

10 papers

cs.RO2026

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

Yajie Li, Bozhou Zhang, Chun Gu +5

Video generation models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively exploiting these imagined fut…

cs.CV2026

See Tomorrow, Act Today: Foresight-Driven Autonomous Driving

Bozhou Zhang, Nan Song, Yuang Wang +3

Current end-to-end autonomous driving planners are fundamentally reactive: they condition on historical and present observations to predict future actions. We argue that autonomous…

cs.RO2026

Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining

Wenyao Zhang, Bozhou Zhang, Zekun Qi +3

Vision-language-action (VLA) models have shown great potential in building generalist robots, but still face a dilemma-misalignment of 2D image forecasting and 3D action prediction…

cs.CV2026

ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving

Jingyu Li, Bozhou Zhang, Xin Jin +3

Autonomous driving requires rich contextual comprehension and precise predictive reasoning to navigate dynamic and complex environments safely. Vision-Language Models (VLMs) and Dr…

cs.CV2025

Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution

Bozhou Zhang, Nan Song, Jingyu Li +3

End-to-end autonomous driving methods aim to directly map raw sensor inputs to future driving actions such as planned trajectories, bypassing traditional modular pipelines. While t…

cs.CV2025

LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving

Nan Song, Bozhou Zhang, Xiatian Zhu +2

Large vision-language models (VLMs) have shown promising capabilities in scene understanding, enhancing the explainability of driving behaviors and interactivity with users. Existi…