5 papers
Learning 4D Geometric Priors for Inference-Efficient World Action Models
Jianjun Zhang, Jian Zhu, Taiyi Su +4
World Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable action sequences. However, existing video-…
PiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation
Chong Ma, Taiyi Su, Jian Zhu +4
Vision-language-action (VLA) policies operate in a closed loop in real-world robot tasks: a robot observes the scene, executes an action chunk, and conditions its next decision on…
Towards High-Consistency Embodied World Model with Multi-View Trajectory Videos
Taiyi Su, Jian Zhu, Yaxuan Li +5
Embodied world models aim to predict and interact with the physical world through visual observations and actions. However, existing models struggle to accurately translate low-lev…
Two-Stream Interactive Joint Learning of Scene Parsing and Geometric Vision Tasks
Guanfeng Tang, Hongbo Zhao, Ziwei Long +5
Inspired by the human visual system, which operates on two parallel yet interactive streams for contextual and spatial understanding, this article presents Two Interactive Streams…
Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving
Tengpeng Li, Hanli Wang, Xianfei Li +3
Autonomous driving is a challenging task that requires perceiving and understanding the surrounding environment for safe trajectory planning. While existing vision-based end-to-end…