6 papers
Conservative Offline Robot Policy Learning via Posterior-Transition Reweighting
Wanpeng Zhang, Hao Luo, Sipeng Zheng +6
Offline post-training adapts a pretrained robot policy to a target dataset by supervised regression on recorded actions. In practice, robot datasets are heterogeneous: they mix emb…
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
Hao Luo, Ye Wang, Wanpeng Zhang +5
Despite progress, Vision-Language-Action models (VLAs) are limited by a scarcity of large-scale, diverse robot data. While human manipulation videos offer a rich alternative, exist…
PlayerOne: Egocentric World Simulator
Yuanpeng Tu, Hao Luo, Xi Chen +3
We introduce PlayerOne, the first egocentric realistic world simulator, facilitating immersive and unrestricted exploration within vividly dynamic environments. Given an egocentric…
LayerFlow: A Unified Model for Layer-aware Video Generation
Sihui Ji, Hao Luo, Xi Chen +3
We present LayerFlow, a unified solution for layer-aware video generation. Given per-layer prompts, LayerFlow generates videos for the transparent foreground, clean background, and…
VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control
Yuanpeng Tu, Hao Luo, Xi Chen +3
Despite significant advancements in video generation, inserting a given object into videos remains a challenging task. The difficulty lies in preserving the appearance details of t…
FashionComposer: Compositional Fashion Image Generation
Sihui Ji, Yiyang Wang, Xi Chen +3
We present FashionComposer for compositional fashion image generation. Unlike previous methods, FashionComposer is highly flexible. It takes multi-modal input (i.e., text prompt, p…