Adding Conditional Control to Text-to-Image Diffusion Models
arXiv:2302.05543
Abstract
We present ControlNet, a neural network architecture to add spatial conditioning controls to large, pretrained text-to-image diffusion models. ControlNet locks the production-ready large diffusion models, and reuses their deep and robust encoding layers pretrained with billions of images as a strong backbone to learn a diverse set of conditional controls. The neural architecture is connected with "zero convolutions" (zero-initialized convolution layers) that progressively grow the parameters from zero and ensure that no harmful noise could affect the finetuning. We test various conditioning controls, eg, edges, depth, segmentation, human pose, etc, with Stable Diffusion, using single or multiple conditions, with or without prompts. We show that the training of ControlNets is robust with small (<50k) and large (>1m) datasets. Extensive results show that ControlNet may facilitate wider applications to control image diffusion models.
Codes and Supplementary Material: https://github.com/lllyasviel/ControlNet
Cited by in corpus (21)
- Towards Small Object Editing: A Benchmark Dataset and A Training-Free Approach
- Diffusion Models, Image Super-Resolution And Everything: A Survey
- PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like Interactions
- Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance Flow
- Generative Artificial Intelligence Meets Synthetic Aperture Radar: A Survey
- 360-Degree Panorama Generation from Few Unregistered NFoV Images
- GANeRF: Leveraging Discriminators to Optimize Neural Radiance Fields
- Painterly Image Harmonization using Diffusion Model
- BlendScape: Enabling End-User Customization of Video-Conferencing Environments through Generative AI
- MemoVis: A GenAI-Powered Tool for Creating Companion Reference Images for 3D Design Feedback
- Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation
- Improving 2D Human Pose Estimation in Rare Camera Views with Synthetic Data
- AI-generated art perceptions with GenFrame -- an image-generating picture frame
- DREAM: Visual Decoding from Reversing Human Visual System
- Reducing Bias in Pre-trained Models by Tuning while Penalizing Change
- P2I-NET: Mapping Camera Pose to Image via Adversarial Learning for New View Synthesis in Real Indoor Environments
- GEMRec: Towards Generative Model Recommendation
- One-shot Unsupervised Domain Adaptation with Personalized Diffusion Models
- Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities
- SeamlessNeRF: Stitching Part NeRFs with Gradient Propagation
- Semantic Generative Augmentations for Few-Shot Counting