DRIT++: Diverse Image-to-Image Translation via Disentangled Representations
arXiv:1905.01270
Abstract
Image-to-image translation aims to learn the mapping between two visual domains. There are two main challenges for this task: 1) lack of aligned training pairs and 2) multiple possible outputs from a single input image. In this work, we present an approach based on disentangled representation for generating diverse outputs without paired training images. To synthesize diverse outputs, we propose to embed images onto two spaces: a domain-invariant content space capturing shared information across domains and a domain-specific attribute space. Our model takes the encoded content features extracted from a given input and attribute vectors sampled from the attribute space to synthesize diverse outputs at test time. To handle unpaired training data, we introduce a cross-cycle consistency loss based on disentangled representations. Qualitative results show that our model can generate diverse and realistic images on a wide range of tasks without paired training data. For quantitative evaluations, we measure realism with user study and Fréchet inception distance, and measure diversity with the perceptual distance metric, Jensen-Shannon divergence, and number of statistically-different bins.
IJCV Journal extension for ECCV 2018 "Diverse Image-to-Image Translation via Disentangled Representations" arXiv:1808.00948. Project Page: http://vllab.ucmerced.edu/hylee/DRIT_pp/ Code: https://github.com/HsinYingLee/DRIT
References in corpus (5)
Cited by in corpus (69)
- Artistic Glyph Image Synthesis via One-Stage Few-Shot Learning
- StarGAN v2: Diverse Image Synthesis for Multiple Domains
- Joint Discriminative and Generative Learning for Person Re-identification
- UNIT-DDPM: UNpaired Image Translation with Denoising Diffusion Probabilistic Models
- High-Resolution Daytime Translation Without Domain Labels
- A Systematic Survey of Regularization and Normalization in GANs
- Mode Seeking Generative Adversarial Networks for Diverse Image Synthesis
- Landmark Assisted CycleGAN for Cartoon Face Generation
- 3D Object Detection and Pose Estimation of Unseen Objects in Color Images with Local Surface Embeddings
- Is Image-to-Image Translation the Panacea for Multimodal Image Registration? A Comparative Study
- Cali-Sketch: Stroke Calibration and Completion for High-Quality Face Image Generation from Human-Like Sketches
- Structural-analogy from a Single Image Pair
- Compatible and Diverse Fashion Image Inpainting
- Image-to-Image Translation: Methods and Applications
- All about Structure: Adapting Structural Information across Domains for Boosting Semantic Segmentation
- Mimicry: Towards the Reproducibility of GAN Research
- GMM-UNIT: Unsupervised Multi-Domain and Multi-Modal Image-to-Image Translation via Attribute Gaussian Mixture Modeling
- IR-GAN: Image Manipulation with Linguistic Instruction by Increment Reasoning
- Dual Contrastive Learning for Unsupervised Image-to-Image Translation
- Bridging the gap between paired and unpaired medical image translation
- DeepFaceEditing: Deep Face Generation and Editing with Disentangled Geometry and Appearance Control
- Contextual colorization and denoising for low-light ultra high resolution sequences
- The Spatially-Correlative Loss for Various Image Translation Tasks
- PI-REC: Progressive Image Reconstruction Network With Edge and Color Domain
- LADN: Local Adversarial Disentangling Network for Facial Makeup and De-Makeup
- Audio-Driven Emotional Video Portraits
- AniGAN: Style-Guided Generative Adversarial Networks for Unsupervised Anime Face Generation
- All-In-One: Facial Expression Transfer, Editing and Recognition Using A Single Network
- Closing the Loop: Joint Rain Generation and Removal via Disentangled Image Translation
- Learning Multi-Site Harmonization of Magnetic Resonance Images Without Traveling Human Phantoms
- Label-Noise Robust Multi-Domain Image-to-Image Translation
- Adversarial Graph Disentanglement
- A Novel BiLevel Paradigm for Image-to-Image Translation
- JOKR: Joint Keypoint Representation for Unsupervised Cross-Domain Motion Retargeting
- Live Face De-Identification in Video
- NAS-DIP: Learning Deep Image Prior with Neural Architecture Search
- Unpaired Image-to-Image Translation using Adversarial Consistency Loss
- Show, Match and Segment: Joint Weakly Supervised Learning of Semantic Matching and Object Co-segmentation
- PREGAN: Pose Randomization and Estimation for Weakly Paired Image Style Translation
- Example-Guided Scene Image Synthesis using Masked Spatial-Channel Attention and Patch-Based Self-Supervision
- Multimodal Image-to-Image Translation via Mutual Information Estimation and Maximization
- FaceShapeGene: A Disentangled Shape Representation for Flexible Face Image Editing
- Image-to-Image Translation with Low Resolution Conditioning
- Im2Pencil: Controllable Pencil Illustration from Photographs
- SDA-GAN: Unsupervised Image Translation Using Spectral Domain Attention-Guided Generative Adversarial Network
- Distribution Aligned Multimodal and Multi-Domain Image Stylization
- Towards Learning a Self-inverse Network for Bidirectional Image-to-image Translation
- One-to-one Mapping for Unpaired Image-to-image Translation
- Attribute-Driven Spontaneous Motion in Unpaired Image Translation
- Face sketch to photo translation using generative adversarial networks
- Reducing Overlearning through Disentangled Representations by Suppressing Unknown Tasks
- CoPE: Conditional image generation using Polynomial Expansions
- DEAAN: Disentangled Embedding and Adversarial Adaptation Network for Robust Speaker Representation Learning
- Memory-guided Unsupervised Image-to-image Translation
- Modeling Artistic Workflows for Image Generation and Editing
- Partially-Shared Variational Auto-encoders for Unsupervised Domain Adaptation with Target Shift
- Unified cross-modality feature disentangler for unsupervised multi-domain MRI abdomen organs segmentation
- Neural Crossbreed: Neural Based Image Metamorphosis
- Smoothing the Disentangled Latent Style Space for Unsupervised Image-to-Image Translation
- Neural Wireframe Renderer: Learning Wireframe to Image Translations
- Semantic View Synthesis
- Multi-scale Neural ODEs for 3D Medical Image Registration
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motion
- Exemplar-Based 3D Portrait Stylization
- A Novel Framework for Image-to-image Translation and Image Compression
- Coarse-to-Fine Gaze Redirection with Numerical and Pictorial Guidance
- Unsupervised Image Transformation Learning via Generative Adversarial Networks
- Delving into Rectifiers in Style-Based Image Translation
- Synthesizing Photorealistic Images with Deep Generative Learning