Disentangling factors of variation in deep representations using adversarial training
arXiv:1611.03383
Abstract
We introduce a conditional generative model for learning to disentangle the hidden factors of variation within a set of labeled observations, and separate them into complementary codes. One code summarizes the specified factors of variation associated with the labels. The other summarizes the remaining unspecified variability. During training, the only available source of supervision comes from our ability to distinguish among different observations belonging to the same class. Examples of such observations include images of a set of labeled objects captured at different viewpoints, or recordings of set of speakers dictating multiple phrases. In both instances, the intra-class diversity is the source of the unspecified factors of variation: each object is observed at multiple viewpoints, and each speaker dictates multiple phrases. Learning to disentangle the specified factors from the unspecified ones becomes easier when strong supervision is possible. Suppose that during training, we have access to pairs of images, where each pair shows two different objects captured from the same viewpoint. This source of alignment allows us to solve our task using existing methods. However, labels for the unspecified factors are usually unavailable in realistic scenarios where data acquisition is not strictly controlled. We address the problem of disentanglement in this more general setting by combining deep convolutional autoencoders with a form of adversarial training. Both factors of variation are implicitly captured in the organization of the learned embedding space, and can be used for solving single-image analogies. Experimental results on synthetic and real datasets show that the proposed method is capable of generalizing to unseen classes and intra-class variabilities.
Conference paper in NIPS 2016
References in corpus (4)
Cited by in corpus (46)
- Recent Advances in Autoencoder-Based Representation Learning
- Fader Networks: Manipulating Images by Sliding Attributes
- Compositional Fairness Constraints for Graph Embeddings
- ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis
- Feature Alignment and Restoration for Domain Generalization and Adaptation
- Finding an Unsupervised Image Segmenter in Each of Your Deep Generative Models
- Learning Interpretable Representation for Controllable Polyphonic Music Generation
- Learning Disentangled Representations with Reference-Based Variational Autoencoders
- Joint Disentangling and Adaptation for Cross-Domain Person Re-Identification
- JADE: Joint Autoencoders for Dis-Entanglement
- Semi-Supervised Generative Modeling for Controllable Speech Synthesis
- Attribute Guided Unpaired Image-to-Image Translation with Semi-supervised Learning
- Enforcing Encoder-Decoder Modularity in Sequence-to-Sequence Models
- StyleMeUp: Towards Style-Agnostic Sketch-Based Image Retrieval
- Learning Cross-domain Generalizable Features by Representation Disentanglement
- Product of Orthogonal Spheres Parameterization for Disentangled Representation Learning
- Transformers with Competitive Ensembles of Independent Mechanisms
- Disentangled Representation Learning with Wasserstein Total Correlation
- Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis
- Adversarially Approximated Autoencoder for Image Generation and Manipulation
- Conditional Adversarial Generative Flow for Controllable Image Synthesis
- Dual Gaussian-based Variational Subspace Disentanglement for Visible-Infrared Person Re-Identification
- DynamicVAE: Decoupling Reconstruction Error and Disentangled Representation Learning
- An Improved Semi-Supervised VAE for Learning Disentangled Representations
- Latent feature disentanglement for 3D meshes
- DualDis: Dual-Branch Disentangling with Adversarial Learning
- Adversarial Disentanglement with Grouped Observations
- Learning Independently-Obtainable Reward Functions
- Unsupervised Shape and Pose Disentanglement for 3D Meshes
- Identity-aware Facial Expression Recognition in Compressed Video
- Generative and Discriminative Learning for Distorted Image Restoration
- Disentangling Dynamics and Content for Control and Planning
- Inference-InfoGAN: Inference Independence via Embedding Orthogonal Basis Expansion
- An Interactive Insight Identification and Annotation Framework for Power Grid Pixel Maps using DenseU-Hierarchical VAE
- StyleFusion: A Generative Model for Disentangling Spatial Segments
- Improving Performance of Seen and Unseen Speech Style Transfer in End-to-end Neural TTS
- Unsupervised Domain Alignment to Mitigate Low Level Dataset Biases
- Domain2Vec: Domain Embedding for Unsupervised Domain Adaptation
- Learning Invariant Representation of Tasks for Robust Surgical State Estimation
- Decoder-free Robustness Disentanglement without (Additional) Supervision
- Towards Purely Unsupervised Disentanglement of Appearance and Shape for Person Images Generation
- Cross-Domain Image Manipulation by Demonstration
- Disentangled Adversarial Transfer Learning for Physiological Biosignals
- Human Annotations Improve GAN Performances
- A Multi-Task Approach for Disentangling Syntax and Semantics in Sentence Representations
- Mutual Information-based Disentangled Neural Networks for Classifying Unseen Categories in Different Domains: Application to Fetal Ultrasound Imaging