Unsupervised Learning of Disentangled Representations from Video
arXiv:1705.10915
Abstract
We present a new model DrNET that learns disentangled image representations from video. Our approach leverages the temporal coherence of video and a novel adversarial loss to learn a representation that factorizes each frame into a stationary part and a temporally varying component. The disentangled representation can be used for a range of tasks. For example, applying a standard LSTM to the time-vary components enables prediction of future frames. We evaluate our approach on a range of synthetic and real videos, demonstrating the ability to coherently generate hundreds of steps into the future.
Cited by in corpus (40)
- Recent Advances in Autoencoder-Based Representation Learning
- A Review on Deep Learning Techniques for Video Prediction
- Multimodal Unsupervised Image-to-Image Translation
- Towards a Definition of Disentangled Representations
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- Stochastic Adversarial Video Prediction
- DRIT++: Diverse Image-to-Image Translation via Disentangled Representations
- Dynamical Variational Autoencoders: A Comprehensive Review
- Video-to-Video Synthesis
- Self-supervised learning of a facial attribute embedding from video
- Learning to Decompose and Disentangle Representations for Video Prediction
- Fashion-Gen: The Generative Fashion Dataset and Challenge
- Symbolic Pregression: Discovering Physical Laws from Distorted Video
- IB-GAN: Disentangled Representation Learning with Information Bottleneck Generative Adversarial Networks
- Mutually improved endoscopic image synthesis and landmark detection in unpaired image-to-image translation
- Adversarially Regularized Autoencoders
- Density-aware Haze Image Synthesis by Self-Supervised Content-Style Disentanglement
- D-GAN: Deep Generative Adversarial Nets for Spatio-Temporal Prediction
- Structural-analogy from a Single Image Pair
- The Sparse Manifold Transform
- Disentangling Factors of Variation by Mixing Them
- Pose Guided Human Video Generation
- Flow-Grounded Spatial-Temporal Video Prediction from Still Images
- Multi-target Voice Conversion without Parallel Data by Adversarially Learning Disentangled Audio Representations
- Few-shot Video-to-Video Synthesis
- TransMoMo: Invariance-Driven Unsupervised Video Motion Retargeting
- Attentive Action and Context Factorization
- Hyperprior Induced Unsupervised Disentanglement of Latent Representations
- Future Frame Prediction for Robot-assisted Surgery
- Learning Controllable Disentangled Representations with Decorrelation Regularization
- Recurrent Flow-Guided Semantic Forecasting
- Disentanglement Learning via Topology
- Motion Selective Prediction for Video Frame Synthesis
- Predicting the Future with Transformational States
- Disentangling Video with Independent Prediction
- Inserting Videos into Videos
- A Deeper Look at the Unsupervised Learning of Disentangled Representations in -VAE from the Perspective of Core Object Recognition
- From Demonstrations to Task-Space Specifications: Using Causal Analysis to Extract Rule Parameterization from Demonstrations
- The pursuit of beauty: Converting image labels to meaningful vectors
- MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics