4 papers
Let Your Video Listen to Your Music!
Xinyu Zhang, Dong Gong, Zicheng Duan +2
Aligning the rhythm of visual motion in a video with a given music track is a practical need in multimedia production, yet remains an underexplored task in autonomous video editing…
EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance
Zicheng Duan, Yuxuan Ding, Chenhui Gou +3
Zero-shot personalized image generation models aim to produce images that align with both a given text prompt and subject image, requiring the model to incorporate both sources of…
HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
Ziqin Zhou, Yifan Yang, Yuqing Yang +7
Text-to-video generation poses significant challenges due to the inherent complexity of video data, which spans both temporal and spatial dimensions. It introduces additional redun…
Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss
Xinyu Zhang, Zicheng Duan, Dong Gong +1
In this paper, we address the challenge of generating temporally consistent videos with motion guidance. While many existing methods depend on additional control modules or inferen…