7 papers
Latent Spatial Memory for Video World Models
Weijie Wang, Haoyu Zhao, Yifan Yang +7
Video world models that maintain 3D spatial consistency across generated frames typically rely on explicit point cloud memory constructed in RGB space. This design is both computat…
Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
Jiatong Xia, Zicheng Duan, Anton van den Hengel +1
Recent progress in 3D generation has been driven largely by models conditioned on images or text, while readily available 3D priors are still underused. In many real-world scenario…
LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
Zicheng Duan, Jiatong Xia, Zeyu Zhang +7
Recent generative video world models aim to simulate visual environment evolution, allowing an observer to interactively explore the scene via camera control. However, they implici…
Frame-wise Conditioning Adaptation for Fine-Tuning Diffusion Models in Text-to-Video Prediction
Zheyuan Liu, Junyan Wang, Zicheng Duan +2
Text-video prediction (TVP) is a downstream video generation task that requires a model to produce subsequent video frames given a series of initial video frames and text describin…
Let Your Video Listen to Your Music!
Xinyu Zhang, Dong Gong, Zicheng Duan +2
Aligning the rhythm of visual motion in a video with a given music track is a practical need in multimedia production, yet remains an underexplored task in autonomous video editing…
EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance
Zicheng Duan, Yuxuan Ding, Chenhui Gou +3
Zero-shot personalized image generation models aim to produce images that align with both a given text prompt and subject image, requiring the model to incorporate both sources of…