papers

Publications (27)

cs.CV2026

FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation

Eric Tillmann Bill, Enis Simsar, Alessio Tonioni +1

Modern text-to-image diffusion models encode rich visual priors, but expose them only through one-way text-conditioned generation. Existing unified vision--language models derived…

cs.CV2026

Shifting the Breaking Point of Flow Matching for Multi-Instance Editing

Carmine Zaccagnino, Fabio Quattrini, Enis Simsar +4

Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offering faster inference through cont…

cs.CV2025

JEDI: The Force of Jensen-Shannon Divergence in Disentangling Diffusion Models

Eric Tillmann Bill, Enis Simsar, Thomas Hofmann

We introduce JEDI, a test-time adaptation method that enhances subject separation and compositional alignment in diffusion models without requiring retraining or external supervisi…

cs.CV2023

LatentSwap3D: Semantic Edits on 3D Image GANs

Enis Simsar, Alessio Tonioni, Evin Pınar Örnek +1

3D GANs have the ability to generate latent codes for entire 3D volumes rather than only 2D images. These models offer desirable features like high-quality geometry and multi-view…

cs.CV2025

UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint

Enis Simsar, Alessio Tonioni, Yongqin Xian +2

We propose an unsupervised instruction-based image editing approach that removes the need for ground-truth edited images during training. Existing methods rely on supervised learni…

cs.LG2026

When One Adapter Speaks for Many: Discovering Low-Rank Redundancy in Continual Fine-Tuning

Tanguy Dieudonné, Giulia Lanzillotta, Enis Simsar +2

Low-Rank Adaptation (LoRA) has become the standard tool for parameter-efficient fine-tuning of large pretrained models. When applied sequentially across tasks in Continual Learning…

cs.CV2026

Editing Everything Everywhere All at Once

Fabio Quattrini, Carmine Zaccagnino, Enis Simsar +4

Editing multiple elements of an image in a single forward pass is a practical alternative to multi-turn image manipulation, offering improved efficiency and potentially better harm…

cs.CV2026

RefAM: Attention Magnets for Zero-Shot Referral Segmentation

Anna Kukleva, Enis Simsar, Alessio Tonioni +4

Most existing approaches to referring segmentation achieve strong performance only through fine-tuning or by composing multiple pre-trained models, often at the cost of additional…

cs.CV2022

Object-aware Monocular Depth Prediction with Instance Convolutions

Enis Simsar, Evin Pınar Örnek, Fabian Manhardt +3

With the advent of deep learning, estimating depth from a single RGB image has recently received a lot of attention, being capable of empowering many different applications ranging…

cs.CV2025

Contrastive Test-Time Composition of Multiple LoRA Models for Image Generation

Tuna Han Salih Meral, Enis Simsar, Federico Tombari +1

Low-Rank Adaptation (LoRA) has emerged as a powerful and popular technique for personalization, enabling efficient adaptation of pre-trained image generation models for specific ta…

cs.CV2023

Diffusion-Based Hierarchical Multi-Label Object Detection to Analyze Panoramic Dental X-rays

Ibrahim Ethem Hamamci, Sezgin Er, Enis Simsar +5

Due to the necessity for precise treatment planning, the use of panoramic X-rays to identify different dental diseases has tremendously increased. Although numerous ML models have…

cs.CV2026

Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography

Ibrahim Ethem Hamamci, Sezgin Er, Chenyu Wang +27

Advancements in medical imaging AI, particularly in 3D imaging, have been limited due to the scarcity of comprehensive datasets. We introduce CT-RATE, a public dataset that pairs 3…

cs.CV2024

GenerateCT: Text-Conditional Generation of 3D Chest CT Volumes

Ibrahim Ethem Hamamci, Sezgin Er, Anjany Sekuboyina +13

GenerateCT, the first approach to generating 3D medical imaging conditioned on free-form medical text prompts, incorporates a text encoder and three key components: a novel causal…

cs.CV2025

Stylebreeder: Exploring and Democratizing Artistic Styles through Text-to-Image Models

Matthew Zheng, Enis Simsar, Hidir Yesiltepe +3

Text-to-image models are becoming increasingly popular, revolutionizing the landscape of digital art creation by enabling highly detailed and creative visual content generation. Th…

cs.CV2022

Fantastic Style Channels and Where to Find Them: A Submodular Framework for Discovering Diverse Directions in GANs

Enis Simsar, Umut Kocasari, Ezgi Gülperi Er +1

The discovery of interpretable directions in the latent spaces of pre-trained GAN models has recently become a popular topic. In particular, StyleGAN2 has enabled various image gen…

cs.CV2023

CONFORM: Contrast is All You Need For High-Fidelity Text-to-Image Diffusion Models

Tuna Han Salih Meral, Enis Simsar, Federico Tombari +1

Images produced by text-to-image diffusion models might not always faithfully represent the semantic intent of the provided text prompt, where the model might overlook or entirely…

cs.CV2024

LIME: Localized Image Editing via Attention Regularization in Diffusion Models

Enis Simsar, Alessio Tonioni, Yongqin Xian +2

Diffusion models (DMs) have gained prominence due to their ability to generate high-quality varied images with recent advancements in text-to-image generation. The research focus i…

cs.CV2025

LoRACLR: Contrastive Adaptation for Customization of Diffusion Models

Enis Simsar, Thomas Hofmann, Federico Tombari +1

Recent advances in text-to-image customization have enabled high-fidelity, context-rich generation of personalized images, allowing specific concepts to appear in a variety of scen…

cs.CV2024

PixLens: A Novel Framework for Disentangled Evaluation in Diffusion-Based Image Editing with Object Detection + SAM

Stefan Stefanache, Lluís Pastor Pérez, Julen Costa Watanabe +3

Evaluating diffusion-based image-editing models is a crucial task in the field of Generative AI. Specifically, it is imperative to assess their capacity to execute diverse editing…

cs.CV2025

IC-Portrait: In-Context Matching for View-Consistent Personalized Portrait

Han Yang, Enis Simsar, Sotiris Anagnostidis +3

Existing diffusion models show great potential for identity-preserving generation. However, personalized portrait generation remains challenging due to the diversity in user profil…

cs.CV2025

SHYI: Action Support for Contrastive Learning in High-Fidelity Text-to-Image Generation

Tianxiang Xia, Lin Xiao, Yannick Montorfani +3

In this project, we address the issue of infidelity in text-to-image generation, particularly for actions involving multiple objects. For this we build on top of the CONFORM framew…

cs.LG2026

SeqLoRA: Bilevel Orthogonal Adaptation for Continual Multi-Concept Generation

Javad Parsa, Enis Simsar, Amir Joudaki +2

Parameter-efficient fine-tuning enables fast personalization of text-to-image diffusion models, but composing multiple custom concepts remains challenging due to representation int…

cs.LG2021

LatentCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions

Oğuz Kaan Yüksel, Enis Simsar, Ezgi Gülperi Er +1

Recent research has shown that it is possible to find interpretable directions in the latent spaces of pre-trained Generative Adversarial Networks (GANs). These directions enable c…

cs.CV2021

Graph2Pix: A Graph-Based Image to Image Translation Framework

Dilara Gokay, Enis Simsar, Efehan Atici +3

In this paper, we propose a graph-based image-to-image translation framework for generating images. We use rich data collected from the popular creativity platform Artbreeder (http…

cs.CV2026

FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation

Eric Tillmann Bill, Enis Simsar, Thomas Hofmann

Text-to-image (T2I) models excel on single-entity prompts but struggle with multi-entity scenes, often exhibiting attribute leakage, identity entanglement, and subject omissions. W…

cs.CV2025

DENTEX: Dental Enumeration and Tooth Pathosis Detection Benchmark for Panoramic X-ray

Ibrahim Ethem Hamamci, Sezgin Er, Omer Faruk Durugol +40

Panoramic X-rays are frequently used in dentistry for treatment planning, but their interpretation can be both time-consuming and prone to error. Artificial intelligence (AI) has t…

cs.CV2024

MegaPortrait: Revisiting Diffusion Control for High-fidelity Portrait Generation

Han Yang, Sotiris Anagnostidis, Enis Simsar +1

We propose MegaPortrait. It's an innovative system for creating personalized portrait images in computer vision. It has three modules: Identity Net, Shading Net, and Harmonization…