Showing cs.ROShow all
3 papers · 1 filter
cs.RO2026
OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies
Yunzhou Song, Long Le, Yong-Hyun Park +7
Vision-language-action(VLA) models have shown great promise as generalist policies for a large range of relatively simple tasks. However, they demonstrate limited performance on mo…
cs.RO2025
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
Junyao Shi, Rujia Yang, Kaitian Chao +9
Today's best-explored routes towards generalist robots center on collecting ever larger "observations-in actions-out" robotics datasets to train large end-to-end models, copying a…
cs.RO2025
Avi: Action from Volumetric Inference
Harris Song, Long Le
We propose Avi, a novel 3D Vision-Language-Action (VLA) architecture that reframes robotic action generation as a problem of 3D perception and spatial reasoning, rather than low-le…