42 papers
Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
Davide Lobba, Fulvio Sanguigni, Bin Ren +3
Recent advances in Virtual Try-On (VTON) and Virtual Try-Off (VTOFF) have greatly improved photo-realistic fashion synthesis and garment reconstruction. However, existing datasets…
Editing Everything Everywhere All at Once
Fabio Quattrini, Carmine Zaccagnino, Enis Simsar +4
Editing multiple elements of an image in a single forward pass is a practical alternative to multi-turn image manipulation, offering improved efficiency and potentially better harm…
A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts
Fabio Quattrini, Carmine Zaccagnino, Costanza Bianchi +2
In this work, we target Handwritten Text Recognition (HTR) in low-resource scenarios, which arise from underrepresented languages, rare scripts, and degraded visual conditions typi…
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
Carmine Zaccagnino, Fabio Quattrini, Enis Simsar +4
Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offering faster inference through cont…
Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation
Tobia Poppi, Silvia Cappelletti, Sara Sarto +5
Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interven…
Hallucination Early Detection in Diffusion Models
Federico Betti, Lorenzo Baraldi, Rita Cucchiara +1
Text-to-Image generation has seen significant advancements in output realism with the advent of diffusion models. However, diffusion models encounter difficulties when tasked with…