Temporal Generative Adversarial Nets with Singular Value Clipping
arXiv:1611.06624
Abstract
In this paper, we propose a generative model, Temporal Generative Adversarial Nets (TGAN), which can learn a semantic representation of unlabeled videos, and is capable of generating videos. Unlike existing Generative Adversarial Nets (GAN)-based methods that generate videos with a single generator consisting of 3D deconvolutional layers, our model exploits two different types of generators: a temporal generator and an image generator. The temporal generator takes a single latent variable as input and outputs a set of latent variables, each of which corresponds to an image frame in a video. The image generator transforms a set of such latent variables into a video. To deal with instability in training of GAN with such advanced networks, we adopt a recently proposed model, Wasserstein GAN, and propose a novel method to train it stably in an end-to-end manner. The experimental results demonstrate the effectiveness of our methods.
to appear in ICCV 2017
References in corpus (10)
- Conditional Generative Adversarial Nets
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Densely Connected Convolutional Networks
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- Improved Techniques for Training GANs
- DRAW: A Recurrent Neural Network For Image Generation
- Generating Videos with Scene Dynamics
- Fully Convolutional Networks for Semantic Segmentation
- Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
Cited by in corpus (10)
- Differentially Private Generative Adversarial Network
- Mode Regularized Generative Adversarial Networks
- Maximum-Likelihood Augmented Discrete Generative Adversarial Networks
- Loss-Sensitive Generative Adversarial Networks on Lipschitz Densities
- Dynamics Transfer GAN: Generating Video by Transferring Arbitrary Temporal Dynamics from a Source Video to a Single Target Image
- Hybrid Learning of Optical Flow and Next Frame Prediction to Boost Optical Flow in the Wild
- To Create What You Tell: Generating Videos from Captions
- Pose Guided Human Video Generation
- Conditional Video Generation Using Action-Appearance Captions
- G3AN: Disentangling Appearance and Motion for Video Generation