5 papers
FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space
FSVideo Team, Qingyu Chen, Zhiyuan Fang +17
We introduce FSVideo, a fast speed transformer-based image-to-video (I2V) diffusion framework. We build our framework on the following key components: 1.) a new video autoencoder w…
ATI: Any Trajectory Instruction for Controllable Video Generation
Angtian Wang, Haibin Huang, Jacob Zhiyuan Fang +2
We propose a unified framework for motion control in video generation that seamlessly integrates camera movement, object-level translation, and fine-grained local motion using traj…
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
Yufan Deng, Yuanyang Yin, Xun Guo +8
We tackle the task of any-reference video generation, which aims to synthesize videos conditioned on arbitrary types and combinations of reference subjects, together with textual p…
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
Yufan Deng, Xun Guo, Yizhi Wang +7
Video generation has witnessed remarkable progress with the advent of deep generative models, particularly diffusion models. While existing methods excel in generating high-quality…
Spatial and Surface Correspondence Field for Interaction Transfer
Zeyu Huang, Honghao Xu, Haibin Huang +3
In this paper, we introduce a new method for the task of interaction transfer. Given an example interaction between a source object and an agent, our method can automatically infer…