7 papers
UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matching
Qilin Huang, Quynh Anh Huynh, Long Le +5
Existing feed-forward networks excel at predicting a single set of physical properties from visual appearance, but this point-estimate paradigm fundamentally fails to capture the r…
OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies
Yunzhou Song, Long Le, Yong-Hyun Park +7
Vision-language-action(VLA) models have shown great promise as generalist policies for a large range of relatively simple tasks. However, they demonstrate limited performance on mo…
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
Junyao Shi, Rujia Yang, Kaitian Chao +9
Today's best-explored routes towards generalist robots center on collecting ever larger "observations-in actions-out" robotics datasets to train large end-to-end models, copying a…
Avi: Action from Volumetric Inference
Harris Song, Long Le
We propose Avi, a novel 3D Vision-Language-Action (VLA) architecture that reframes robotic action generation as a problem of 3D perception and spatial reasoning, rather than low-le…
Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels
Long Le, Ryan Lucas, Chen Wang +4
Inferring the physical properties of 3D scenes from visual information is a critical yet challenging task for creating interactive and realistic virtual worlds. While humans intuit…
Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model
Long Le, Jason Xie, William Liang +7
Interactive 3D simulated objects are crucial in AR/VR, animations, and robotics, driving immersive experiences and advanced automation. However, creating these articulated objects…