collaborators
Showing cs.ROShow all

8 papers · 1 filter

cs.RO2026

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

AgiBot Research Team, Renhang Liu, Wenzhi Zhao +42

World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video…

cs.RO2026

-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

Xiaowei Cai, Yunuo Cai, Bingao Chen +36

Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-acti…

cs.RO2026

-WM: A Unified Video-Action World Model for Robotic Manipulation

Pengfei Zhou, Shengcong Chen, Di Chen +17

Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present -World…

cs.RO2025

Act2Goal: From World Model To General Goal-conditioned Policy

Pengfei Zhou, Liliang Chen, Shengcong Chen +5

Specifying robotic manipulation tasks in a manner that is both expressive and precise remains a central challenge. While visual goals provide a compact and unambiguous task specifi…

cs.RO2025

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Yue Liao, Pengfei Zhou, Siyuan Huang +11

We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that integrates policy learning, evaluation, and simulation within a single video-g…

cs.RO2025

EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models

Hu Yue, Siyuan Huang, Yue Liao +5

Recent advances in creative AI have enabled the synthesis of high-fidelity images and videos conditioned on language instructions. Building on these developments, text-to-video dif…