11 papers
Physical Object Understanding with a Physically Controllable World Model
Rahul Venkatesh, Klemen Kotar, Lilian Naing Chen +9
A central challenge in visual intelligence is learning the physical structure of scenes from raw videos: how regions form objects and the laws that govern their interactions. Solvi…
GE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic Manipulation
Boxiang Qiu, Liliang Chen, Yue Liao +12
We introduce GE-Sim 2.0 (Genie Envisioner World Simulator 2.0), a closed-loop video world simulator for robotic manipulation. Building on the action-conditioned video generation fr…
Zero-shot World Models Are Developmentally Efficient Learners
Khai Loong Aw, Klemen Kotar, Wanhee Lee +6
Young children demonstrate early abilities to understand their physical world, estimating depth, motion, object coherence, interactions, and many other aspects of physical scene un…
Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface
Yujie Zhao, Hongwei Fan, Di Chen +5
Recent progress in robot learning has been driven by large-scale datasets and powerful visuomotor policy architectures, yet policy robustness remains limited by the substantial cos…
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
Yi Liu, Sukai Wang, Dafeng Wei +10
General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execution, a combination that remains challeng…
Act2Goal: From World Model To General Goal-conditioned Policy
Pengfei Zhou, Liliang Chen, Shengcong Chen +5
Specifying robotic manipulation tasks in a manner that is both expressive and precise remains a central challenge. While visual goals provide a compact and unambiguous task specifi…