17 papers
EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
Jiayi Luo, Hanxin Zhu, Chen Gao +5
Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable p…
CP4D: Compositional Physics-aware 4D Scene Generation
Hanxin Zhu, Cong Wang, Tianyu He +4
4D generation (\textit{i.e.}, dynamic 3D generation) has recently emerged as a rapidly growing research frontier due to its powerful spatiotemporal modeling capabilities. However,…
Beyond Pixel Histories: World Models with Persistent 3D State
Samuel Garcin, Thomas Walker, Steven McDonagh +5
Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities. However, existing models typically lack a 3D rep…
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
Yang Yue, Fangyun Wei, Tianyu He +10
Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregressive generators built on d…
GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion
Hanxin Zhu, Cong Wang, Peiyan Tu +4
Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of domains including spatial intellige…
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
Haoyu Wu, Diankun Wu, Tianyu He +4
Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw video data often fail to capture…