3 papers
cs.RO2025
VLA-0: Building State-of-the-Art VLAs with Zero Modification
Ankit Goyal, Hugo Hadfield, Xuning Yang +2
Vision-Language-Action models (VLAs) hold immense promise for enabling generalist robot manipulation. However, the best way to build them remains an open question. Current approach…
cs.GR2025
3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds
Fan-Yun Sun, Shengguang Wu, Christian Jacobsen +13
Despite large-scale pretraining endowing models with language and vision reasoning capabilities, improving their spatial reasoning capability remains challenging due to the lack of…
cs.RO2025
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
Ishika Singh, Ankit Goyal, Stan Birchfield +3
We introduce OG-VLA, a novel architecture and learning framework that combines the generalization strengths of Vision Language Action models (VLAs) with the robustness of 3D-aware…