ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation
arXiv:2311.18822 · doi:10.1109/cvpr52733.2024.00631
Abstract
Diffusion models have revolutionized image generation in recent years, yet they are still limited to a few sizes and aspect ratios. We propose ElasticDiffusion, a novel training-free decoding method that enables pretrained text-to-image diffusion models to generate images with various sizes. ElasticDiffusion attempts to decouple the generation trajectory of a pretrained model into local and global signals. The local signal controls low-level pixel information and can be estimated on local patches, while the global signal is used to maintain overall structural consistency and is estimated with a reference image. We test our method on CelebA-HQ (faces) and LAION-COCO (objects/indoor/outdoor scenes). Our experiments and qualitative results show superior image coherence quality across aspect ratios compared to MultiDiffusion and the standard decoding strategy of Stable Diffusion. Project page: https://elasticdiffusion.github.io/
Accepted at CVPR 2024. Project Page: https://elasticdiffusion.github.io/
References in corpus (32)
- LoRA: Low-Rank Adaptation of Large Language Models
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Diffusion Models Beat GANs on Image Synthesis
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- LAION-5B: An open large-scale dataset for training next generation image-text models
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Blended Diffusion for Text-driven Editing of Natural Images
- Classifier-Free Diffusion Guidance
- Cascaded Diffusion Models for High Fidelity Image Generation
- Prompt-to-Prompt Image Editing with Cross Attention Control
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Diffusion Posterior Sampling for General Noisy Inverse Problems
- MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation
- Novel View Synthesis with Diffusion Models
- On Fast Sampling of Diffusion Probabilistic Models
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models
- Latent Video Diffusion Models for High-Fidelity Long Video Generation
- Inst-Inpaint: Instructing to Remove Objects with Diffusion Models
- ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion Models
- Mixture of Diffusers for scene composition and high resolution image generation
- AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation
- Training-free Diffusion Model Adaptation for Variable-Sized Text-to-Image Synthesis
- AnimateLCM: Computation-Efficient Personalized Style Video Generation without Personalized Video Data
- Portrait Diffusion: Training-free Face Stylization with Chain-of-Painting
- Fast High-Resolution Image Synthesis with Latent Adversarial Diffusion Distillation
- LoCo: Locally Constrained Training-Free Layout-to-Image Synthesis
- Curved Diffusion: A Generative Model With Optical Geometry Control
- MagicScroll: Nontypical Aspect-Ratio Image Generation for Visual Storytelling via Multi-Layered Semantic-Aware Denoising
- Adaptive Guidance: Training-free Acceleration of Conditional Diffusion Models
- HiDiffusion: Unlocking Higher-Resolution Creativity and Efficiency in Pretrained Diffusion Models
- Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation