10 papers
Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation
Shubhankar Borse, Phuc Pham, Farzad Farhadzadeh +6
Despite recent advances in personalized image generation, existing models consistently fail to produce reliable multi-human scenes, often merging or losing facial identity. We pres…
Resolving the Identity Crisis in Text-to-Image Generation
Shubhankar Borse, Farzad Farhadzadeh, Munawar Hayat +1
State-of-the-art text-to-image models suffer from a persistent identity crisis when generating scenes with multiple humans: producing duplicate faces, merging identities, and misco…
MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
Shubhankar Borse, Seokeon Choi, Sunghyun Park +6
Generation of images containing multiple humans, performing complex actions, while preserving their facial identities, is a significant challenge. A major factor contributing to th…
MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
Shreya Kadambi, Risheek Garrepalli, Shubhankar Borse +2
Despite the remarkable success of diffusion models in text-to-image generation, their effectiveness in grounded visual editing and compositional control remains challenging. Motiva…
Personalized OVSS: Understanding Personal Concept in Open-Vocabulary Semantic Segmentation
Sunghyun Park, Jungsoo Lee, Shubhankar Borse +4
While open-vocabulary semantic segmentation (OVSS) can segment an image into semantic regions based on arbitrarily given text descriptions even for classes unseen during training,…
Zero-Shot Adaptation of Parameter-Efficient Fine-Tuning in Diffusion Models
Farzad Farhadzadeh, Debasmit Das, Shubhankar Borse +1
We introduce ProLoRA, enabling zero-shot adaptation of parameter-efficient fine-tuning in text-to-image diffusion models. ProLoRA transfers pre-trained low-rank adjustments (e.g.,…