MotionVideoGAN: A Novel Video Generator Based on the Motion Space Learned from Image Pairs
arXiv:2303.02906 · doi:10.1109/TMM.2023.3251095
Abstract
Video generation has achieved rapid progress benefiting from high-quality renderings provided by powerful image generators. We regard the video synthesis task as generating a sequence of images sharing the same contents but varying in motions. However, most previous video synthesis frameworks based on pre-trained image generators treat content and motion generation separately, leading to unrealistic generated videos. Therefore, we design a novel framework to build the motion space, aiming to achieve content consistency and fast convergence for video generation. We present MotionVideoGAN, a novel video generator synthesizing videos based on the motion space learned by pre-trained image pair generators. Firstly, we propose an image pair generator named MotionStyleGAN to generate image pairs sharing the same contents and producing various motions. Then we manage to acquire motion codes to edit one image in the generated image pairs and keep the other unchanged. The motion codes help us edit images within the motion space since the edited image shares the same contents with the other unchanged one in image pairs. Finally, we introduce a latent code generator to produce latent code sequences using motion codes for video generation. Our approach achieves state-of-the-art performance on the most complex video dataset ever used for unconditional video generation evaluation, UCF101.
Accepted by IEEE Transactions on Multimedia as a regular paper
References in corpus (6)
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Alias-Free Generative Adversarial Networks
- Generating Videos with Scene Dynamics
- StyleNeRF: A Style-based 3D-Aware Generator for High-resolution Image Synthesis
- Controlling generative models with continuous factors of variations
- You Only Need Adversarial Supervision for Semantic Image Synthesis