3 papers
cs.RO2026
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
Zaibin Zhang, Junlan Xiao, Zhongbo Zhang +11
Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) models integrate perception, language, and control, but most rep…
cs.CV2025
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
Zhida Zhao, Talas Fu, Yifan Wang +2
Despite remarkable progress in driving world models, their potential for autonomous systems remains largely untapped: the world models are mostly learned for world simulation and d…
cs.CV2024
AD-H: Language-guided Autonomous Driving with Hierarchical Agents
Zaibin Zhang, Talas Fu, Shiyu Tang +4
Language-guided autonomous driving requires bridging a large abstraction gap between high-level natural-language instructions and low-level vehicle control. End-to-end approaches t…