3 papers
cs.RO2026
DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation
Jian Zhu, Jianjun Zhang, Taiyi Su +10
World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learni…
cs.CV2026
ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
Yuhao Zhou, Yunpeng Zhu, Yang Zhou +7
Vision-Language-Action (VLA) models hold great promise for general-purpose robotic intelligence, yet scaling up such models is severely bottlenecked by the high cost of acquiring a…
cs.RO2025
UnderwaterVLA: Dual-brain Vision-Language-Action architecture for Autonomous Underwater Navigation
Zhangyuan Wang, Yunpeng Zhu, Yuqi Yan +7
This paper presents UnderwaterVLA, a novel framework for autonomous underwater navigation that integrates multimodal foundation models with embodied intelligence systems. Underwate…