5 papers · 1 filter
Do-Undo Bench: Reversibility for Action Understanding in Image Generation
Shweta Mahajan, Shreya Kadambi, Hoang Le +4
We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transformations driven by real-world a…
MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
Shubhankar Borse, Seokeon Choi, Sunghyun Park +6
Generation of images containing multiple humans, performing complex actions, while preserving their facial identities, is a significant challenge. A major factor contributing to th…
MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
Shreya Kadambi, Risheek Garrepalli, Shubhankar Borse +2
Despite the remarkable success of diffusion models in text-to-image generation, their effectiveness in grounded visual editing and compositional control remains challenging. Motiva…
SubZero: Composing Subject, Style, and Action via Zero-Shot Personalization
Shubhankar Borse, Kartikeya Bhardwaj, Mohammad Reza Karimi Dastjerdi +8
Diffusion models are increasingly popular for generative tasks, including personalized composition of subjects and styles. While diffusion models can generate user-specified subjec…
FouRA: Fourier Low Rank Adaptation
Shubhankar Borse, Shreya Kadambi, Nilesh Prasad Pandey +7
While Low-Rank Adaptation (LoRA) has proven beneficial for efficiently fine-tuning large models, LoRA fine-tuned text-to-image diffusion models lack diversity in the generated imag…