6 papers
PlayWorld: Learning Robot World Models from Autonomous Play
Tenny Yin, Zhiting Mei, Zhonghe Zheng +8
Action-conditioned video models offer a promising path to building general-purpose robot simulators that can improve directly from data. Yet, despite training on large-scale robot…
World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
Zhiting Mei, Tenny Yin, Micah Baker +2
Recent advances in generative video models have led to significant breakthroughs in high-fidelity video synthesis, specifically in controllable video generation where the generated…
LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
Lihan Zha, Asher J. Hancock, Mingtong Zhang +5
A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodim…
Video Generation Models in Robotics -- Applications, Research Challenges, Future Directions
Zhiting Mei, Tenny Yin, Ola Shorinwa +9
Video generation models have emerged as high-fidelity models of the physical world, capable of synthesizing high-quality videos capturing fine-grained interactions between agents a…
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
Zhiting Mei, Christina Zhang, Tenny Yin +3
Reasoning language models have set state-of-the-art (SOTA) records on many challenging benchmarks, enabled by multi-step reasoning induced using reinforcement learning. However, li…
WoMAP: World Models For Embodied Open-Vocabulary Object Localization
Tenny Yin, Zhiting Mei, Tao Sun +6
Language-instructed active object localization is a critical challenge for robots, requiring efficient exploration of partially observable environments. However, state-of-the-art a…