Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation
arXiv:2407.03006 · doi:10.1609/aaai.v38i3.27951
Abstract
Recently, large-scale text-to-image (T2I) diffusion models have emerged as a powerful tool for image-to-image translation (I2I), allowing open-domain image translation via user-provided text prompts. This paper proposes frequency-controlled diffusion model (FCDiffusion), an end-to-end diffusion-based framework that contributes a novel solution to text-guided I2I from a frequency-domain perspective. At the heart of our framework is a feature-space frequency-domain filtering module based on Discrete Cosine Transform, which filters the latent features of the source image in the DCT domain, yielding filtered image features bearing different DCT spectral bands as different control signals to the pre-trained Latent Diffusion Model. We reveal that control signals of different DCT spectral bands bridge the source image and the T2I generated image in different correlations (e.g., style, structure, layout, contour, etc.), and thus enable versatile I2I applications emphasizing different I2I correlations, including style-guided content creation, image semantic manipulation, image scene translation, and image style translation. Different from related approaches, FCDiffusion establishes a unified text-guided I2I framework suitable for diverse image translation tasks simply by switching among different frequency control branches at inference time. The effectiveness and superiority of our method for text-guided I2I are demonstrated with extensive experiments both qualitatively and quantitatively. Our project is publicly available at: https://xianggao1102.github.io/FCDiffusion/.
Proceedings of the 38th AAAI Conference on Artificial Intelligence (AAAI 2024)
References in corpus (26)
- Denoising Diffusion Probabilistic Models
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks
- Diffusion Models Beat GANs on Image Synthesis
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Prompt-to-Prompt Image Editing with Cross Attention Control
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
- Adding Conditional Control to Text-to-Image Diffusion Models
- Contrastive Learning for Unpaired Image-to-Image Translation
- StarGAN v2: Diverse Image Synthesis for Multiple Domains
- Pretraining is All You Need for Image-to-Image Translation
- RePaint: Inpainting using Denoising Diffusion Probabilistic Models
- Diffusion-based Image Translation using Disentangled Style and Content Representation
- InstructPix2Pix: Learning to Follow Image Editing Instructions
- DiffusionCLIP: Text-Guided Diffusion Models for Robust Image Manipulation
- VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance
- Text2LIVE: Text-Driven Layered Image and Video Editing
- Diffusion Probabilistic Models for 3D Point Cloud Generation
- Geometry-Consistent Generative Adversarial Networks for One-Sided Unsupervised Domain Mapping
- RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and Generation
- Frequency Domain Image Translation: More Photo-realistic, Better Identity-preserving
- Splicing ViT Features for Semantic Appearance Transfer
- Learning Frequency-aware Dynamic Network for Efficient Super-Resolution
- Learning to Incorporate Texture Saliency Adaptive Attention to Image Cartoonization
- DCT-Conv: Coding filters in convolutional networks with Discrete Cosine Transform