3 papers
cs.RO2026
DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models
Emily Yue-Ting Jia, Weiduo Yuan, Tianheng Shi +3
Robotic manipulation requires sophisticated commonsense reasoning, a capability naturally possessed by large-scale Vision-Language Models (VLMs). While VLMs show promise as zero-sh…
cs.RO2025
Robot Learning from Any Images
Siheng Zhao, Jiageng Mao, Wei Chow +11
We introduce RoLA, a framework that transforms any in-the-wild image into an interactive, physics-enabled robotic environment. Unlike previous methods, RoLA operates directly on a…
cs.RO2024
Learning from Massive Human Videos for Universal Humanoid Pose Control
Jiageng Mao, Siheng Zhao, Siqi Song +7
Scalable learning of humanoid robots is crucial for their deployment in real-world applications. While traditional approaches primarily rely on reinforcement learning or teleoperat…