most citedMoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation

1 citations · 2 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2025

Preventing Shortcuts in Adapter Training via Providing the Shortcuts

Anujraaj Argo Goyal, Guocheng Gordon Qian, Huseyin Coskun +8

Adapter-based training has emerged as a key mechanism for extending the capabilities of powerful foundation image generators, enabling personalized and stylized text-to-image synth…

cs.CV2025

ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation

Guocheng Gordon Qian, Daniil Ostashev, Egor Nemchinov +4

Generating high-fidelity images of humans with fine-grained control over attributes such as hairstyle and clothing remains a core challenge in personalized text-to-image synthesis.…

cs.CV2025

CanvasComposer: Personalized Group Photo Generation via a Multi-Reference Canvas

Gordon Guocheng Qian, Ruihang Zhang, Tsai-Shien Chen +11

Existing personalized image generators still struggle to preserve multiple reference identities in natural and coherent multi-human generations. To address these limitations, we pr…

cs.CV2025

Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing

Rishubh Parihar, Or Patashnik, Daniil Ostashev +3

Instruction-based image editing offers a powerful and intuitive way to manipulate images through natural language. Yet, relying solely on text instructions limits fine-grained cont…

cs.CV2025

Scaling Group Inference for Diverse and High-Quality Generation

Gaurav Parmar, Or Patashnik, Daniil Ostashev +4

Generative models typically sample outputs independently, and recent inference-time guidance and scaling algorithms focus on improving the quality of individual samples. However, i…

cs.CV20251 cited

Object-level Visual Prompts for Compositional Image Generation

Gaurav Parmar, Or Patashnik, Kuan-Chieh Wang +5

We introduce a method for composing object-level visual prompts within a text-to-image diffusion model. Our approach addresses the task of generating semantically coherent composit…