16 papers
Editing Everything Everywhere All at Once
Fabio Quattrini, Carmine Zaccagnino, Enis Simsar +4
Editing multiple elements of an image in a single forward pass is a practical alternative to multi-turn image manipulation, offering improved efficiency and potentially better harm…
When One Adapter Speaks for Many: Discovering Low-Rank Redundancy in Continual Fine-Tuning
Tanguy Dieudonné, Giulia Lanzillotta, Enis Simsar +2
Low-Rank Adaptation (LoRA) has become the standard tool for parameter-efficient fine-tuning of large pretrained models. When applied sequentially across tasks in Continual Learning…
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
Carmine Zaccagnino, Fabio Quattrini, Enis Simsar +4
Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offering faster inference through cont…
SeqLoRA: Bilevel Orthogonal Adaptation for Continual Multi-Concept Generation
Javad Parsa, Enis Simsar, Amir Joudaki +2
Parameter-efficient fine-tuning enables fast personalization of text-to-image diffusion models, but composing multiple custom concepts remains challenging due to representation int…
FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation
Eric Tillmann Bill, Enis Simsar, Alessio Tonioni +1
Modern text-to-image diffusion models encode rich visual priors, but expose them only through one-way text-conditioned generation. Existing unified vision--language models derived…
FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
Eric Tillmann Bill, Enis Simsar, Thomas Hofmann
Text-to-image (T2I) models excel on single-entity prompts but struggle with multi-entity scenes, often exhibiting attribute leakage, identity entanglement, and subject omissions. W…