Deep Convolutional Inverse Graphics Network
arXiv:1503.03167
Abstract
This paper presents the Deep Convolution Inverse Graphics Network (DC-IGN), a model that learns an interpretable representation of images. This representation is disentangled with respect to transformations such as out-of-plane rotations and lighting variations. The DC-IGN model is composed of multiple layers of convolution and de-convolution operators and is trained using the Stochastic Gradient Variational Bayes (SGVB) algorithm. We propose a training procedure to encourage neurons in the graphics code layer to represent a specific transformation (e.g. pose or light). Given a single input image, our model can generate new images of the same object with variations in pose and lighting. We present qualitative and quantitative results of the model's efficacy at learning a 3D rendering engine.
First two authors contributed equally
References in corpus (3)
Cited by in corpus (50)
- Variational Autoencoder for Deep Learning of Images, Labels and Captions
- Recent Advances in Autoencoder-Based Representation Learning
- Fader Networks: Manipulating Images by Sliding Attributes
- Disentangling factors of variation in deep representations using adversarial training
- Learning What and Where to Draw
- MONet: Unsupervised Scene Decomposition and Representation
- Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data
- Learning to Generate Images of Outdoor Scenes from Attributes and Semantic Layouts
- On the Origin of Deep Learning
- DeepStereo: Learning to Predict New Views from the World's Imagery
- Video Frame Interpolation via Adaptive Separable Convolution
- Deferred Neural Rendering: Image Synthesis using Neural Textures
- GeneGAN: Learning Object Transfiguration and Attribute Subspace from Unpaired Data
- Convolutional Network for Attribute-driven and Identity-preserving Human Face Generation
- Tackling Over-pruning in Variational Autoencoders
- Emergence of Exploratory Look-Around Behaviors through Active Observation Completion
- Transformation-Grounded Image Generation Network for Novel 3D View Synthesis
- A Survey of Inductive Biases for Factorial Representation-Learning
- Video Frame Interpolation via Adaptive Convolution
- Adversarial examples for generative models
- Unsupervised Visual Attribute Transfer with Reconfigurable Generative Adversarial Networks
- A Survey on Deep Learning Architectures for Image-based Depth Reconstruction
- Relevance Factor VAE: Learning and Identifying Disentangled Factors
- Learning Disentangled Representations with Reference-Based Variational Autoencoders
- 3D Shape Induction from 2D Views of Multiple Objects
- Disentangling Space and Time in Video with Hierarchical Variational Auto-encoders
- JADE: Joint Autoencoders for Dis-Entanglement
- DeepWarp: Photorealistic Image Resynthesis for Gaze Manipulation
- Inducing Interpretable Representations with Variational Autoencoders
- Label Denoising Adversarial Network (LDAN) for Inverse Lighting of Face Images
- Material Editing Using a Physically Based Rendering Network
- Unpaired Pose Guided Human Image Generation
- Deep Variational Inference Without Pixel-Wise Reconstruction
- Cerberus: A Multi-headed Derenderer
- Attentive Action and Context Factorization
- Learning Disentangled Representations of Timbre and Pitch for Musical Instrument Sounds Using Gaussian Mixture Variational Autoencoders
- Latent feature disentanglement for 3D meshes
- An Uncertain Future: Forecasting from Static Images using Variational Autoencoders
- CDVAE: Co-embedding Deep Variational Auto Encoder for Conditional Variational Generation
- Soft Constraints for Inference with Declarative Knowledge
- Affine Variational Autoencoders: An Efficient Approach for Improving Generalization and Robustness to Distribution Shift
- Improving Bi-directional Generation between Different Modalities with Variational Autoencoders
- Deep Structure for end-to-end inverse rendering
- Thinking Outside the Pool: Active Training Image Creation for Relative Attributes
- Inferring 3D Shapes from Image Collections using Adversarial Networks
- Disentangling Video with Independent Prediction
- Learning to Recognize Objects by Retaining other Factors of Variation
- Learning to Look Around: Intelligently Exploring Unseen Environments for Unknown Tasks
- Deep disentangled representations for volumetric reconstruction
- Attribute-controlled face photo synthesis from simple line drawing