7 papers
FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
Eric Tillmann Bill, Enis Simsar, Thomas Hofmann
Text-to-image (T2I) models excel on single-entity prompts but struggle with multi-entity scenes, often exhibiting attribute leakage, identity entanglement, and subject omissions. W…
LoRACLR: Contrastive Adaptation for Customization of Diffusion Models
Enis Simsar, Thomas Hofmann, Federico Tombari +1
Recent advances in text-to-image customization have enabled high-fidelity, context-rich generation of personalized images, allowing specific concepts to appear in a variety of scen…
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
Enis Simsar, Alessio Tonioni, Yongqin Xian +2
We propose an unsupervised instruction-based image editing approach that removes the need for ground-truth edited images during training. Existing methods rely on supervised learni…
JEDI: The Force of Jensen-Shannon Divergence in Disentangling Diffusion Models
Eric Tillmann Bill, Enis Simsar, Thomas Hofmann
We introduce JEDI, a test-time adaptation method that enhances subject separation and compositional alignment in diffusion models without requiring retraining or external supervisi…
IC-Portrait: In-Context Matching for View-Consistent Personalized Portrait
Han Yang, Enis Simsar, Sotiris Anagnostidis +3
Existing diffusion models show great potential for identity-preserving generation. However, personalized portrait generation remains challenging due to the diversity in user profil…
SHYI: Action Support for Contrastive Learning in High-Fidelity Text-to-Image Generation
Tianxiang Xia, Lin Xiao, Yannick Montorfani +3
In this project, we address the issue of infidelity in text-to-image generation, particularly for actions involving multiple objects. For this we build on top of the CONFORM framew…