3 papers
cs.RO2025
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
Junyao Shi, Rujia Yang, Kaitian Chao +9
Today's best-explored routes towards generalist robots center on collecting ever larger "observations-in actions-out" robotics datasets to train large end-to-end models, copying a…
cs.RO2025
Avi: Action from Volumetric Inference
Harris Song, Long Le
We propose Avi, a novel 3D Vision-Language-Action (VLA) architecture that reframes robotic action generation as a problem of 3D perception and spatial reasoning, rather than low-le…
cs.CV2025
Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels
Long Le, Ryan Lucas, Chen Wang +4
Inferring the physical properties of 3D scenes from visual information is a critical yet challenging task for creating interactive and realistic virtual worlds. While humans intuit…