15 papers
EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
Jiayi Luo, Hanxin Zhu, Chen Gao +5
Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable p…
CoT-Edit: Let CoT Guide Instruction Video Editing
Sen Liang, Fengbin Guan, Youliang Zhang +2
Text-driven instruction-based video editing in complex scenes remains challenging: purely textual prompts often fail to capture precise spatial relationships and physical constrain…
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
Sen Liang, Cong Wang, Zhentao Yu +8
Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet the complex creative demands of real-world scenarios. To bridge…
DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning
Zili Lin, Wenyao Zhang, Yuyang Zhang +9
Demonstration augmentation is proposed for cost-efficient data acquisition, but existing methods are fundamentally limited in deformable manipulation due to two challenges: (1) the…
Embody4D: A Generalist Data Engine for Embodied 4D World Modeling
Peiyan Tu, Hanxin Zhu, Jingwen Sun +6
Embodied agents require robust and comprehensive 3D spatiotemporal representations to support spatial reasoning, manipulation understanding, and downstream decision making. However…
Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment
Cong Wang, Hanxin Zhu, Jiayi Luo +6
Large-scale video generation models have made remarkable progress in semantic consistency and visual quality, producing videos that are increasingly coherent and visually convincin…