activity
20242026
collaborators

16 papers

cs.RO2026

EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation

Jiayi Luo, Hanxin Zhu, Chen Gao +5

Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable p…

cs.RO2026

DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning

Zili Lin, Wenyao Zhang, Yuyang Zhang +9

Demonstration augmentation is proposed for cost-efficient data acquisition, but existing methods are fundamentally limited in deformable manipulation due to two challenges: (1) the…

cs.CV2026

CP4D: Compositional Physics-aware 4D Scene Generation

Hanxin Zhu, Cong Wang, Tianyu He +4

4D generation (\textit{i.e.}, dynamic 3D generation) has recently emerged as a rapidly growing research frontier due to its powerful spatiotemporal modeling capabilities. However,…

cs.CV2026

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling

Peiyan Tu, Hanxin Zhu, Jingwen Sun +6

Embodied agents require robust and comprehensive 3D spatiotemporal representations to support spatial reasoning, manipulation understanding, and downstream decision making. However…

cs.CV2026

Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment

Cong Wang, Hanxin Zhu, Jiayi Luo +6

Large-scale video generation models have made remarkable progress in semantic consistency and visual quality, producing videos that are increasingly coherent and visually convincin…

cs.CV2026

OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance

Cong Wang, Hanxin Zhu, Xiao Tang +4

Recent progress in video generation has led to substantial improvements in visual fidelity, yet ensuring physically consistent motion remains a fundamental challenge. Intuitively,…