Publications (27)
FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation
Eric Tillmann Bill, Enis Simsar, Alessio Tonioni +1
Modern text-to-image diffusion models encode rich visual priors, but expose them only through one-way text-conditioned generation. Existing unified vision--language models derived…
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
Carmine Zaccagnino, Fabio Quattrini, Enis Simsar +4
Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offering faster inference through cont…
JEDI: The Force of Jensen-Shannon Divergence in Disentangling Diffusion Models
Eric Tillmann Bill, Enis Simsar, Thomas Hofmann
We introduce JEDI, a test-time adaptation method that enhances subject separation and compositional alignment in diffusion models without requiring retraining or external supervisi…
LatentSwap3D: Semantic Edits on 3D Image GANs
Enis Simsar, Alessio Tonioni, Evin Pınar Ãrnek +1
3D GANs have the ability to generate latent codes for entire 3D volumes rather than only 2D images. These models offer desirable features like high-quality geometry and multi-view…
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
Enis Simsar, Alessio Tonioni, Yongqin Xian +2
We propose an unsupervised instruction-based image editing approach that removes the need for ground-truth edited images during training. Existing methods rely on supervised learni…
When One Adapter Speaks for Many: Discovering Low-Rank Redundancy in Continual Fine-Tuning
Tanguy Dieudonné, Giulia Lanzillotta, Enis Simsar +2
Low-Rank Adaptation (LoRA) has become the standard tool for parameter-efficient fine-tuning of large pretrained models. When applied sequentially across tasks in Continual Learning…
Editing Everything Everywhere All at Once
Fabio Quattrini, Carmine Zaccagnino, Enis Simsar +4
Editing multiple elements of an image in a single forward pass is a practical alternative to multi-turn image manipulation, offering improved efficiency and potentially better harm…
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
Anna Kukleva, Enis Simsar, Alessio Tonioni +4
Most existing approaches to referring segmentation achieve strong performance only through fine-tuning or by composing multiple pre-trained models, often at the cost of additional…
Object-aware Monocular Depth Prediction with Instance Convolutions
Enis Simsar, Evin Pınar Ãrnek, Fabian Manhardt +3
With the advent of deep learning, estimating depth from a single RGB image has recently received a lot of attention, being capable of empowering many different applications ranging…
Contrastive Test-Time Composition of Multiple LoRA Models for Image Generation
Tuna Han Salih Meral, Enis Simsar, Federico Tombari +1
Low-Rank Adaptation (LoRA) has emerged as a powerful and popular technique for personalization, enabling efficient adaptation of pre-trained image generation models for specific ta…
Diffusion-Based Hierarchical Multi-Label Object Detection to Analyze Panoramic Dental X-rays
Ibrahim Ethem Hamamci, Sezgin Er, Enis Simsar +5
Due to the necessity for precise treatment planning, the use of panoramic X-rays to identify different dental diseases has tremendously increased. Although numerous ML models have…
Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
Ibrahim Ethem Hamamci, Sezgin Er, Chenyu Wang +27
Advancements in medical imaging AI, particularly in 3D imaging, have been limited due to the scarcity of comprehensive datasets. We introduce CT-RATE, a public dataset that pairs 3…
GenerateCT: Text-Conditional Generation of 3D Chest CT Volumes
Ibrahim Ethem Hamamci, Sezgin Er, Anjany Sekuboyina +13
GenerateCT, the first approach to generating 3D medical imaging conditioned on free-form medical text prompts, incorporates a text encoder and three key components: a novel causal…
Stylebreeder: Exploring and Democratizing Artistic Styles through Text-to-Image Models
Matthew Zheng, Enis Simsar, Hidir Yesiltepe +3
Text-to-image models are becoming increasingly popular, revolutionizing the landscape of digital art creation by enabling highly detailed and creative visual content generation. Th…
Fantastic Style Channels and Where to Find Them: A Submodular Framework for Discovering Diverse Directions in GANs
Enis Simsar, Umut Kocasari, Ezgi Gülperi Er +1
The discovery of interpretable directions in the latent spaces of pre-trained GAN models has recently become a popular topic. In particular, StyleGAN2 has enabled various image gen…
CONFORM: Contrast is All You Need For High-Fidelity Text-to-Image Diffusion Models
Tuna Han Salih Meral, Enis Simsar, Federico Tombari +1
Images produced by text-to-image diffusion models might not always faithfully represent the semantic intent of the provided text prompt, where the model might overlook or entirely…
LIME: Localized Image Editing via Attention Regularization in Diffusion Models
Enis Simsar, Alessio Tonioni, Yongqin Xian +2
Diffusion models (DMs) have gained prominence due to their ability to generate high-quality varied images with recent advancements in text-to-image generation. The research focus i…
LoRACLR: Contrastive Adaptation for Customization of Diffusion Models
Enis Simsar, Thomas Hofmann, Federico Tombari +1
Recent advances in text-to-image customization have enabled high-fidelity, context-rich generation of personalized images, allowing specific concepts to appear in a variety of scen…
PixLens: A Novel Framework for Disentangled Evaluation in Diffusion-Based Image Editing with Object Detection + SAM
Stefan Stefanache, LluÃs Pastor Pérez, Julen Costa Watanabe +3
Evaluating diffusion-based image-editing models is a crucial task in the field of Generative AI. Specifically, it is imperative to assess their capacity to execute diverse editing…
IC-Portrait: In-Context Matching for View-Consistent Personalized Portrait
Han Yang, Enis Simsar, Sotiris Anagnostidis +3
Existing diffusion models show great potential for identity-preserving generation. However, personalized portrait generation remains challenging due to the diversity in user profil…
SHYI: Action Support for Contrastive Learning in High-Fidelity Text-to-Image Generation
Tianxiang Xia, Lin Xiao, Yannick Montorfani +3
In this project, we address the issue of infidelity in text-to-image generation, particularly for actions involving multiple objects. For this we build on top of the CONFORM framew…
SeqLoRA: Bilevel Orthogonal Adaptation for Continual Multi-Concept Generation
Javad Parsa, Enis Simsar, Amir Joudaki +2
Parameter-efficient fine-tuning enables fast personalization of text-to-image diffusion models, but composing multiple custom concepts remains challenging due to representation int…
LatentCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions
OÄuz Kaan Yüksel, Enis Simsar, Ezgi Gülperi Er +1
Recent research has shown that it is possible to find interpretable directions in the latent spaces of pre-trained Generative Adversarial Networks (GANs). These directions enable c…
Graph2Pix: A Graph-Based Image to Image Translation Framework
Dilara Gokay, Enis Simsar, Efehan Atici +3
In this paper, we propose a graph-based image-to-image translation framework for generating images. We use rich data collected from the popular creativity platform Artbreeder (http…
FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
Eric Tillmann Bill, Enis Simsar, Thomas Hofmann
Text-to-image (T2I) models excel on single-entity prompts but struggle with multi-entity scenes, often exhibiting attribute leakage, identity entanglement, and subject omissions. W…
DENTEX: Dental Enumeration and Tooth Pathosis Detection Benchmark for Panoramic X-ray
Ibrahim Ethem Hamamci, Sezgin Er, Omer Faruk Durugol +40
Panoramic X-rays are frequently used in dentistry for treatment planning, but their interpretation can be both time-consuming and prone to error. Artificial intelligence (AI) has t…
MegaPortrait: Revisiting Diffusion Control for High-fidelity Portrait Generation
Han Yang, Sotiris Anagnostidis, Enis Simsar +1
We propose MegaPortrait. It's an innovative system for creating personalized portrait images in computer vision. It has three modules: Identity Net, Shading Net, and Harmonization…