Blended Latent Diffusion
arXiv:2206.02779 · doi:10.1145/3592450
Abstract
The tremendous progress in neural image generation, coupled with the emergence of seemingly omnipotent vision-language models has finally enabled text-based interfaces for creating and editing images. Handling generic images requires a diverse underlying generative model, hence the latest works utilize diffusion models, which were shown to surpass GANs in terms of diversity. One major drawback of diffusion models, however, is their relatively slow inference time. In this paper, we present an accelerated solution to the task of local text-driven editing of generic images, where the desired edits are confined to a user-provided mask. Our solution leverages a recent text-to-image Latent Diffusion Model (LDM), which speeds up diffusion by operating in a lower-dimensional latent space. We first convert the LDM into a local image editor by incorporating Blended Diffusion into it. Next we propose an optimization-based solution for the inherent inability of this LDM to accurately reconstruct images. Finally, we address the scenario of performing local edits using thin masks. We evaluate our method against the available baselines both qualitatively and quantitatively and demonstrate that in addition to being faster, our method achieves better precision than the baselines while mitigating some of their artifacts.
Accepted to SIGGRAPH 2023. Project page: https://omriavrahami.com/blended-latent-diffusion-page/
References in corpus (1)
Cited by in corpus (22)
- Diffusion Models in Vision: A Survey
- Towards Small Object Editing: A Benchmark Dataset and A Training-Free Approach
- SpaText: Spatio-Textual Representation for Controllable Image Generation
- Automated data processing and feature engineering for deep learning and big data applications: a survey
- Break-A-Scene: Extracting Multiple Concepts from a Single Image
- Face Generation and Editing with StyleGAN: A Survey
- Diffusion Model-Based Image Editing: A Survey
- Uni-paint: A Unified Framework for Multimodal Image Inpainting with Pretrained Diffusion Model
- The Chosen One: Consistent Characters in Text-to-Image Diffusion Models
- Latent Denoising Diffusion GAN: Faster sampling, Higher image quality
- TexSliders: Diffusion-Based Texture Editing in CLIP Space
- Stable Flow: Vital Layers for Training-Free Image Editing
- Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields
- RadEdit: stress-testing biomedical vision models via diffusion image editing
- DiffUHaul: A Training-Free Method for Object Dragging in Images
- Polyp-Gen: Realistic and Diverse Polyp Image Generation for Endoscopic Dataset Expansion
- Face Aging via Diffusion-based Editing
- PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models
- 3D-Fixup: Advancing Photo Editing with 3D Priors
- From Visual Explanations to Counterfactual Explanations with Latent Diffusion
- ESCT3D: Efficient and Selectively Controllable Text-Driven 3D Content Generation with Gaussian Splatting
- DiffAugment: Diffusion based Long-Tailed Visual Relationship Recognition