18 papers
InSight: Self-Guided Skill Acquisition via Steerable VLAs
Maggie Wang, Lars Osterberg, Stephen Tian +3
Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. We present InSight, a…
PlayWorld: Learning Robot World Models from Autonomous Play
Tenny Yin, Zhiting Mei, Zhonghe Zheng +8
Action-conditioned video models offer a promising path to building general-purpose robot simulators that can improve directly from data. Yet, despite training on large-scale robot…
World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
Zhiting Mei, Tenny Yin, Micah Baker +2
Recent advances in generative video models have led to significant breakthroughs in high-fidelity video synthesis, specifically in controllable video generation where the generated…
Phys2Real: Fusing VLM Priors with Interactive Online Adaptation for Uncertainty-Aware Sim-to-Real Manipulation
Maggie Wang, Stephen Tian, Aiden Swann +3
Learning robotic manipulation policies directly in the real world can be expensive and time-consuming. While reinforcement learning (RL) policies trained in simulation present a sc…
Video Generation Models in Robotics -- Applications, Research Challenges, Future Directions
Zhiting Mei, Tenny Yin, Ola Shorinwa +9
Video generation models have emerged as high-fidelity models of the physical world, capable of synthesizing high-quality videos capturing fine-grained interactions between agents a…
Geometry Meets Vision: Revisiting Pretrained Semantics in Distilled Fields
Zhiting Mei, Ola Shorinwa, Anirudha Majumdar
Semantic distillation in radiance fields has spurred significant advances in open-vocabulary robot policies, e.g., in manipulation and navigation, founded on pretrained semantics f…