Transformation-based Adversarial Video Prediction on Large-Scale Data
arXiv:2003.04035
Abstract
Recent breakthroughs in adversarial generative modeling have led to models capable of producing video samples of high quality, even on large and complex datasets of real-world video. In this work, we focus on the task of video prediction, where given a sequence of frames extracted from a video, the goal is to generate a plausible future sequence. We first improve the state of the art by performing a systematic empirical study of discriminator decompositions and proposing an architecture that yields faster convergence and higher performance than previous approaches. We then analyze recurrent units in the generator, and propose a novel recurrent unit which transforms its past hidden state according to predicted motion-like features, and refines it to handle dis-occlusions, scene changes and other complex behavior. We show that this recurrent unit consistently outperforms previous designs. Our final model leads to a leap in the state-of-the-art performance, obtaining a test set Frechet Video Distance of 25.7, down from 69.2, on the large-scale Kinetics-600 dataset.
References in corpus (13)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- The Kinetics Human Action Video Dataset
- A Learned Representation For Artistic Style
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- Decomposing Motion and Content for Natural Video Sequence Prediction
- A Short Note on the Kinetics-700-2020 Human Action Dataset
- Recurrent Convolutional Strategies for Face Manipulation Detection in Videos
- Modulating early visual processing by language
- Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks
- Geometric GAN
- Deep Visual Foresight for Planning Robot Motion