Separating Content and Style for Unsupervised Image-to-Image Translation
arXiv:2110.14404
Abstract
Unsupervised image-to-image translation aims to learn the mapping between two visual domains with unpaired samples. Existing works focus on disentangling domain-invariant content code and domain-specific style code individually for multimodal purposes. However, less attention has been paid to interpreting and manipulating the translated image. In this paper, we propose to separate the content code and style code simultaneously in a unified framework. Based on the correlation between the latent features and the high-level domain-invariant tasks, the proposed framework demonstrates superior performance in multimodal translation, interpretability and manipulation of the translated image. Experimental results show that the proposed approach outperforms the existing unsupervised image translation methods in terms of visual quality and diversity.
Accepted by BMVC2021
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- CyCADA: Cycle-Consistent Adversarial Domain Adaptation
- MISO: Mutual Information Loss with Stochastic Style Representations for Multimodal Image-to-Image Translation
- TransGaGa: Geometry-Aware Unsupervised Image-to-Image Translation
- Generative Adversarial Network with Multi-Branch Discriminator for Cross-Species Image-to-Image Translation
- Controlling biases and diversity in diverse image-to-image translation
- Improving Style-Content Disentanglement in Image-to-Image Translation