4 papers
DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation
Jian Zhu, Jianjun Zhang, Taiyi Su +10
World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learni…
ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
Yuhao Zhou, Yunpeng Zhu, Yang Zhou +7
Vision-Language-Action (VLA) models hold great promise for general-purpose robotic intelligence, yet scaling up such models is severely bottlenecked by the high cost of acquiring a…
UnderwaterVLA: Dual-brain Vision-Language-Action architecture for Autonomous Underwater Navigation
Zhangyuan Wang, Yunpeng Zhu, Yuqi Yan +7
This paper presents UnderwaterVLA, a novel framework for autonomous underwater navigation that integrates multimodal foundation models with embodied intelligence systems. Underwate…
Scaling User Modeling: Large-scale Online User Representations for Ads Personalization in Meta
Wei Zhang, Dai Li, Chen Liang +17
Effective user representations are pivotal in personalized advertising. However, stringent constraints on training throughput, serving latency, and memory, often limit the complexi…