collaborators

7 papers

cs.CV2026

Latent Spatial Memory for Video World Models

Weijie Wang, Haoyu Zhao, Yifan Yang +7

Video world models that maintain 3D spatial consistency across generated frames typically rely on explicit point cloud memory constructed in RGB space. This design is both computat…

cs.CV2026

Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors

Jiatong Xia, Zicheng Duan, Anton van den Hengel +1

Recent progress in 3D generation has been driven largely by models conditioned on images or text, while readily available 3D priors are still underused. In many real-world scenario…

cs.CV2026

LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models

Zicheng Duan, Jiatong Xia, Zeyu Zhang +7

Recent generative video world models aim to simulate visual environment evolution, allowing an observer to interactively explore the scene via camera control. However, they implici…

cs.CV2025

Frame-wise Conditioning Adaptation for Fine-Tuning Diffusion Models in Text-to-Video Prediction

Zheyuan Liu, Junyan Wang, Zicheng Duan +2

Text-video prediction (TVP) is a downstream video generation task that requires a model to produce subsequent video frames given a series of initial video frames and text describin…

cs.CV2025

Let Your Video Listen to Your Music!

Xinyu Zhang, Dong Gong, Zicheng Duan +2

Aligning the rhythm of visual motion in a video with a given music track is a practical need in multimedia production, yet remains an underexplored task in autonomous video editing…

cs.CV2025

EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance

Zicheng Duan, Yuxuan Ding, Chenhui Gou +3

Zero-shot personalized image generation models aim to produce images that align with both a given text prompt and subject image, requiring the model to incorporate both sources of…