Image-to-image translation for cross-domain disentanglement
arXiv:1805.09730
Abstract
Deep image translation methods have recently shown excellent results, outputting high-quality images covering multiple modes of the data distribution. There has also been increased interest in disentangling the internal representations learned by deep methods to further improve their performance and achieve a finer control. In this paper, we bridge these two objectives and introduce the concept of cross-domain disentanglement. We aim to separate the internal representation into three parts. The shared part contains information for both domains. The exclusive parts, on the other hand, contain only factors of variation that are particular to each domain. We achieve this through bidirectional image translation based on Generative Adversarial Networks and cross-domain autoencoders, a novel network component. Our model offers multiple advantages. We can output diverse samples covering multiple modes of the distributions of both domains, perform domain-specific image transfer and interpolation, and cross-domain retrieval without the need of labeled data, only paired images. We compare our model to the state-of-the-art in multi-modal image translation and achieve better results for translation on challenging datasets as well as for cross-domain retrieval on realistic datasets.
Accepted to NIPS 2018
References in corpus (8)
- InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets
- Toward Multimodal Image-to-Image Translation
- Adversarial Feature Learning
- Adversarially Learned Inference
- Domain Separation Networks
- Augmented CycleGAN: Learning Many-to-Many Mappings from Unpaired Data
- Multimodal Unsupervised Image-to-Image Translation
- Mix and match networks: encoder-decoder alignment for zero-pair image translation
Cited by in corpus (11)
- Doodle to Search: Practical Zero-Shot Sketch-based Image Retrieval
- Artistic Glyph Image Synthesis via One-Stage Few-Shot Learning
- Illumination-Adaptive Person Re-identification
- Contrastive Variational Autoencoder Enhances Salient Features
- Deep Sketch-guided Cartoon Video Inbetweening
- A Generative Adversarial Network for AI-Aided Chair Design
- FMODetect: Robust Detection of Fast Moving Objects
- Towards Disentangled Representations for Human Retargeting by Multi-view Learning
- Learning Disentangled Representations of Satellite Image Time Series
- Disentangling Pose from Appearance in Monochrome Hand Images
- Learning to adapt class-specific features across domains for semantic segmentation