Blended Diffusion for Text-driven Editing of Natural Images
arXiv:2111.14818 · doi:10.1109/CVPR52688.2022.01767
Abstract
Natural language offers a highly intuitive interface for image editing. In this paper, we introduce the first solution for performing local (region-based) edits in generic natural images, based on a natural language description along with an ROI mask. We achieve our goal by leveraging and combining a pretrained language-image model (CLIP), to steer the edit towards a user-provided text prompt, with a denoising diffusion probabilistic model (DDPM) to generate natural-looking results. To seamlessly fuse the edited region with the unchanged parts of the image, we spatially blend noised versions of the input image with the local text-guided diffusion latent at a progression of noise levels. In addition, we show that adding augmentations to the diffusion process mitigates adversarial results. We compare against several baselines and related methods, both qualitatively and quantitatively, and show that our method outperforms these solutions in terms of overall realism, ability to preserve the background and matching the text. Finally, we show several text-driven editing applications, including adding a new object to an image, removing/replacing/altering existing objects, background replacement, and image extrapolation. Code is available at: https://omriavrahami.com/blended-diffusion-page/
CVPR 2022. Code is available at: https://omriavrahami.com/blended-diffusion-page/
References in corpus (7)
- Explaining and Harnessing Adversarial Examples
- Diffusion Models Beat GANs on Image Synthesis
- Zero-Shot Text-to-Image Generation
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
- Alias-Free Generative Adversarial Networks
- More Control for Free! Image Synthesis with Semantic Diffusion Guidance
- Towards Open-World Text-Guided Face Image Generation and Manipulation
Cited by in corpus (37)
- Diffusion Models in Vision: A Survey
- Blended Latent Diffusion
- Towards Small Object Editing: A Benchmark Dataset and A Training-Free Approach
- SpaText: Spatio-Textual Representation for Controllable Image Generation
- Break-A-Scene: Extracting Multiple Concepts from a Single Image
- DiffEdit: Diffusion-based semantic image editing with mask guidance
- Diffusion Model-Based Image Editing: A Survey
- PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like Interactions
- Uni-paint: A Unified Framework for Multimodal Image Inpainting with Pretrained Diffusion Model
- Diffusion Models Meet Remote Sensing: Principles, Methods, and Perspectives
- DiLightNet: Fine-grained Lighting Control for Diffusion-based Image Generation
- FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution
- Sparse Visual Counterfactual Explanations in Image Space
- The Chosen One: Consistent Characters in Text-to-Image Diffusion Models
- Explaining generative diffusion models via visual analysis for interpretable decision-making process
- Iterative Motion Editing with Natural Language
- TexSliders: Diffusion-Based Texture Editing in CLIP Space
- Towards Interactive Image Inpainting via Sketch Refinement
- Stable Flow: Vital Layers for Training-Free Image Editing
- Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields
- RadEdit: stress-testing biomedical vision models via diffusion image editing
- ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation
- DiffUHaul: A Training-Free Method for Object Dragging in Images
- Geometric-Facilitated Denoising Diffusion Model for 3D Molecule Generation
- Magic Fixup: Streamlining Photo Editing by Watching Dynamic Videos
- IntrinsicEdit: Precise generative image manipulation in intrinsic space
- Content-Based Search for Deep Generative Models
- GaussEdit: Adaptive 3D Scene Editing with Text and Image Prompts
- Continual Test-Time Adaptation for Single Image Defocus Deblurring via Causal Siamese Networks
- Face Aging via Diffusion-based Editing
- 3D-Fixup: Advancing Photo Editing with 3D Priors
- TALE: Training-free Cross-domain Image Composition via Adaptive Latent Manipulation and Energy-guided Optimization
- Uncertainty for SVBRDF Acquisition using Frequency Analysis
- Pro-DG: Procedural Diffusion Guidance for Architectural Facade Generation
- Text-Conditioned Background Generation for Editable Multi-Layer Documents
- Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications
- SAT3D: Image-driven Semantic Attribute Transfer in 3D