Weakly-supervised Disentangling with Recurrent Transformations for 3D View Synthesis
arXiv:1601.00706
Abstract
An important problem for both graphics and vision is to synthesize novel views of a 3D object from a single image. This is particularly challenging due to the partial observability inherent in projecting a 3D object onto the image space, and the ill-posedness of inferring object shape and pose. However, we can train a neural network to address the problem if we restrict our attention to specific object categories (in our case faces and chairs) for which we can gather ample training data. In this paper, we propose a novel recurrent convolutional encoder-decoder network that is trained end-to-end on the task of rendering rotated objects starting from a single image. The recurrent structure allows our model to capture long-term dependencies along a sequence of transformations. We demonstrate the quality of its predictions for human faces on the Multi-PIE dataset and for a dataset of 3D chair models, and also show its ability to disentangle latent factors of variation (e.g., identity and pose) without using full supervision.
This was published in NIPS 2015 conference
References in corpus (3)
Cited by in corpus (85)
- Generative Adversarial Text to Image Synthesis
- Deep Face Recognition: A Survey
- Learning-Based View Synthesis for Light Field Cameras
- Pose Guided Person Image Generation
- Stacked Hourglass Networks for Human Pose Estimation
- Dynamic Filter Networks
- Perceptual Adversarial Networks for Image-to-Image Transformation
- Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding
- A Survey on Deep Learning Techniques for Stereo-based Depth Estimation
- Fader Networks: Manipulating Images by Sliding Attributes
- Learning Factorized Multimodal Representations
- Triple Generative Adversarial Nets
- Representation Learning by Rotating Your Faces
- DARLA: Improving Zero-Shot Transfer in Reinforcement Learning
- Beyond Face Rotation: Global and Local Perception GAN for Photorealistic and Identity Preserving Frontal View Synthesis
- Early Visual Concept Learning with Unsupervised Deep Learning
- Are Disentangled Representations Helpful for Abstract Visual Reasoning?
- Weakly-Supervised Disentanglement Without Compromises
- Video Frame Interpolation via Adaptive Separable Convolution
- RenderNet: A deep convolutional network for differentiable rendering from 3D shapes
- Learning Disentangled Joint Continuous and Discrete Representations
- Multiple-Attribute Text Style Transfer
- Variational Inference of Disentangled Latent Concepts from Unlabeled Observations
- Disentangling Factors of Variation Using Few Labels
- Weakly-Supervised Discovery of Geometry-Aware Representation for 3D Human Pose Estimation
- Towards Large-Pose Face Frontalization in the Wild
- Transformation-Grounded Image Generation Network for Novel 3D View Synthesis
- Neural Face Editing with Intrinsic Image Disentangling
- A Recurrent Encoder-Decoder Network for Sequential Face Alignment
- Video Frame Interpolation via Adaptive Convolution
- Challenges in Disentangling Independent Factors of Variation
- Semantically Decomposing the Latent Spaces of Generative Adversarial Networks
- 3D Shape Reconstruction from Sketches via Multi-view Convolutional Networks
- Generative Adversarial Talking Head: Bringing Portraits to Life with a Weakly Supervised Neural Network
- Pose-Invariant Face Alignment with a Single CNN
- View Synthesis by Appearance Flow
- A Survey on Deep Learning Architectures for Image-based Depth Reconstruction
- Multi-Channel Attention Selection GAN with Cascaded Semantic Guidance for Cross-View Image Translation
- Relevance Factor VAE: Learning and Identifying Disentangled Factors
- Load Balanced GANs for Multi-view Face Image Synthesis
- Learning Disentangled Representations with Reference-Based Variational Autoencoders
- Facial UV Map Completion for Pose-invariant Face Recognition: A Novel Adversarial Approach based on Coupled Attention Residual UNets
- Learning to Synthesize a 4D RGBD Light Field from a Single Image
- Ranking CGANs: Subjective Control over Semantic Image Attributes
- Pixels, voxels, and views: A study of shape representations for single view 3D object shape prediction
- Towards Automatic Image Editing: Learning to See another You
- Triple Generative Adversarial Networks
- Transformable Bottleneck Networks
- Disentangled Representation Learning with Wasserstein Total Correlation
- Disentangled Representations in Neural Models
- On the Fairness of Disentangled Representations
- Learning 6-DOF Grasping Interaction via Deep Geometry-aware 3D Representations
- Generative Model with Coordinate Metric Learning for Object Recognition Based on 3D Models
- SpVOS: Efficient Video Object Segmentation with Triple Sparse Convolution
- Visuomotor Understanding for Representation Learning of Driving Scenes
- Measuring the Biases and Effectiveness of Content-Style Disentanglement
- A Neural Rendering Framework for Free-Viewpoint Relighting
- View Extrapolation of Human Body from a Single Image
- DynamicVAE: Decoupling Reconstruction Error and Disentangled Representation Learning
- Latent feature disentanglement for 3D meshes
- Gotta Adapt 'Em All: Joint Pixel and Feature-Level Domain Adaptation for Recognition in the Wild
- Adversarial Disentanglement with Grouped Observations
- Future Urban Scenes Generation Through Vehicles Synthesis
- Hierarchical Transfer Convolutional Neural Networks for Image Classification
- Deep Structure for end-to-end inverse rendering
- Curriculum Learning for Recurrent Video Object Segmentation
- PCLs: Geometry-aware Neural Reconstruction of 3D Pose with Perspective Crop Layers
- Cross-View Image Synthesis with Deformable Convolution and Attention Mechanism
- Recurrent Deconvolutional Generative Adversarial Networks with Application to Text Guided Video Generation
- PI-GAN: Learning Pose Independent representations for multiple profile face synthesis
- Learning to Recognize Objects by Retaining other Factors of Variation
- Deep disentangled representations for volumetric reconstruction
- Profile to Frontal Face Recognition in the Wild Using Coupled Conditional GAN
- 3D Ken Burns Effect from a Single Image
- Quantised Transforming Auto-Encoders: Achieving Equivariance to Arbitrary Transformations in Deep Networks
- A Shape-Aware Retargeting Approach to Transfer Human Motion and Appearance in Monocular Videos
- Novel View Synthesis via Depth-guided Skip Connections
- Learning Equivariant Representations
- An Interpretable Generative Model for Handwritten Digit Image Synthesis
- MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics
- Group-based Learning of Disentangled Representations with Generalizability for Novel Contents
- Robust 2D/3D Vehicle Parsing in CVIS
- Depth Assisted Full Resolution Network for Single Image-based View Synthesis
- Sparse Pose Trajectory Completion
- Inner Space Preserving Generative Pose Machine