7 citations · 11 across the 18 of their papers we have counts for
20 papers · 1 filter
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
Marco Mistretta, Alberto Baldrati, Lorenzo Agnolucci +2
Pre-trained multi-modal Vision-Language Models like CLIP are widely used off-the-shelf for a variety of applications. In this paper, we show that the common practice of individuall…
Garment Attribute Manipulation with Multi-level Attention
Vittorio Casula, Lorenzo Berlincioni, Luca Cultrera +5
In the rapidly evolving field of online fashion shopping, the need for more personalized and interactive image retrieval systems has become paramount. Existing methods often strugg…
Improving Zero-shot Generalization of Learned Prompts via Unsupervised Knowledge Distillation
Marco Mistretta, Alberto Baldrati, Marco Bertini +1
Vision-Language Models (VLMs) demonstrate remarkable zero-shot generalization to unseen tasks, but fall short of the performance of supervised methods in generalizing to downstream…
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
Emanuele Vivoli, Irene Campaioli, Mariateresa Nardoni +3
Comics, as a medium, uniquely combine text and images in styles often distinct from real-world visuals. For the past three decades, computational research on comics has evolved fro…
Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
Alberto Baldrati, Davide Morelli, Marcella Cornia +2
Fashion illustration is a crucial medium for designers to convey their creative vision and transform design concepts into tangible representations that showcase the interplay betwe…
Perceptual Quality Improvement in Videoconferencing using Keyframes-based GAN
Lorenzo Agnolucci, Leonardo Galteri, Marco Bertini +1
In the latest years, videoconferencing has taken a fundamental role in interpersonal relations, both for personal and business purposes. Lossy video compression algorithms are the…