4 papers
GeoCo-SAVi: Geometry-Consistent Slot Attention for Explicitly Editable Object Representations
Haoxiang Huang, Zhekai Wang, Xiang Liu +2
Object-centric video models represent scenes with slots, yet exposed geometry can vary in meaning with appearance. In Invariant Slot Attention (ISA), explicit position and scale ca…
ConfAL-WM: Confidence-Guided Active Learning for Action-Conditioned World Models
Xiang Liu, Sen Cui, Changshui Zhang
Action-conditioned world models have become an important foundation for embodied prediction, planning, and synthetic data generation, but their errors under new task and scene dist…
MOSH-WM: Mask-Grounded Soft-Hamiltonian Dynamics for Object-Centric World Models
Zhekai Wang, Haoxiang Huang, Xiang Liu +6
Object-centric world models forecast future videos by evolving a set of entity slots, but the variables receiving dynamics supervision are often unconstrained visual features. We i…
Scene2Demo: Self-Evolving Embodied Data Generation via Object-Action Graph
Xiang Liu, Sen Cui, Guocai Yao +4
We present Scene2Demo, a self-evolving framework for offline embodied data generation. Given a single real-world RGB image and a user query, Scene2Demo constructs an interactive si…