11 papers
FieldGen: From Teleoperated Pre-Manipulation Trajectories to Field-Guided Data Generation
Wenhao Wang, Kehe Ye, Xinyu Zhou +9
Large-scale and diverse datasets are vital for training robust robotic manipulation policies, yet existing data collection methods struggle to balance scale, diversity, and quality…
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
Yue Liao, Pengfei Zhou, Siyuan Huang +11
We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that integrates policy learning, evaluation, and simulation within a single video-g…
EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models
Hu Yue, Siyuan Huang, Yue Liao +5
Recent advances in creative AI have enabled the synthesis of high-fidelity images and videos conditioned on language instructions. Building on these developments, text-to-video dif…
EnerVerse-AC: Envisioning Embodied Environments with Action Condition
Yuxin Jiang, Shengcong Chen, Siyuan Huang +8
Robotic imitation learning has advanced from solving static tasks to addressing dynamic interaction scenarios, but testing and evaluation remain costly and challenging due to the n…
Hume: Introducing System-2 Thinking in Visual-Language-Action Model
Haoming Song, Delin Qu, Yuanqi Yao +9
Humans practice slow thinking before performing actual actions when handling complex tasks in the physical world. This thinking paradigm, recently, has achieved remarkable advancem…
Genie Centurion: Accelerating Scalable Real-World Robot Training with Human Rewind-and-Refine Guidance
Wenhao Wang, Jianheng Song, Chiming Liu +13
While Vision-Language-Action (VLA) models show strong generalizability in various tasks, real-world deployment of robotic policy still requires large-scale, high-quality human expe…