3 papers
cs.CV2026
Embody4D: A Generalist Data Engine for Embodied 4D World Modeling
Peiyan Tu, Hanxin Zhu, Jingwen Sun +6
Embodied agents require robust and comprehensive 3D spatiotemporal representations to support spatial reasoning, manipulation understanding, and downstream decision making. However…
cs.CV2026
Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment
Cong Wang, Hanxin Zhu, Jiayi Luo +6
Large-scale video generation models have made remarkable progress in semantic consistency and visual quality, producing videos that are increasingly coherent and visually convincin…
cs.CV2026
GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion
Hanxin Zhu, Cong Wang, Peiyan Tu +4
Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of domains including spatial intellige…