15 papers · 1 filter
Continuous Control of Editing Models via Adaptive-Origin Guidance
Alon Wolf, Chen Katzir, Kfir Aberman +1
Diffusion-based editing models have emerged as a powerful tool for semantic image and video manipulation. However, existing models lack a mechanism for smoothly controlling the int…
Canvas-to-Image: Compositional Image Generation with Multimodal Controls
Yusuf Dalva, Guocheng Gordon Qian, Maya Goldenberg +5
While modern diffusion models excel at generating high-quality and diverse images, they still struggle with high-fidelity compositional and multimodal control, particularly when us…
Preventing Shortcuts in Adapter Training via Providing the Shortcuts
Anujraaj Argo Goyal, Guocheng Gordon Qian, Huseyin Coskun +8
Adapter-based training has emerged as a key mechanism for extending the capabilities of powerful foundation image generators, enabling personalized and stylized text-to-image synth…
ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation
Guocheng Gordon Qian, Daniil Ostashev, Egor Nemchinov +4
Generating high-fidelity images of humans with fine-grained control over attributes such as hairstyle and clothing remains a core challenge in personalized text-to-image synthesis.…
Scaling Group Inference for Diverse and High-Quality Generation
Gaurav Parmar, Or Patashnik, Daniil Ostashev +4
Generative models typically sample outputs independently, and recent inference-time guidance and scaling algorithms focus on improving the quality of individual samples. However, i…
Be Decisive: Noise-Induced Layouts for Multi-Subject Generation
Omer Dahary, Yehonathan Cohen, Or Patashnik +2
Generating multiple distinct subjects remains a challenge for existing text-to-image diffusion models. Complex prompts often lead to subject leakage, causing inaccuracies in quanti…