7 citations · 8 across the 7 of their papers we have counts for
7 papers
Improving Zero-shot Generalization of Learned Prompts via Unsupervised Knowledge Distillation
Marco Mistretta, Alberto Baldrati, Marco Bertini +1
Vision-Language Models (VLMs) demonstrate remarkable zero-shot generalization to unseen tasks, but fall short of the performance of supervised methods in generalizing to downstream…
Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
Alberto Baldrati, Davide Morelli, Marcella Cornia +2
Fashion illustration is a crucial medium for designers to convey their creative vision and transform design concepts into tangible representations that showcase the interplay betwe…
Mapping Memes to Words for Multimodal Hateful Meme Classification
Giovanni Burbi, Alberto Baldrati, Lorenzo Agnolucci +2
Multimodal image-text memes are prevalent on the internet, serving as a unique form of communication that combines visual and textual elements to convey humor, ideas, or emotions.…
Exploiting CLIP-based Multi-modal Approach for Artwork Classification and Retrieval
Alberto Baldrati, Marco Bertini, Tiberio Uricchio +1
Given the recent advances in multimodal image pretraining where visual models trained with semantically dense textual supervision tend to have better generalization capabilities th…
OpenFashionCLIP: Vision-and-Language Contrastive Learning with Open-Source Fashion Data
Giuseppe Cartella, Alberto Baldrati, Davide Morelli +3
The inexorable growth of online shopping and e-commerce demands scalable and robust machine learning-based solutions to accommodate customer requirements. In the context of automat…
Composed Image Retrieval using Contrastive Learning and Task-oriented CLIP-based Features
Alberto Baldrati, Marco Bertini, Tiberio Uricchio +1
Given a query composed of a reference image and a relative caption, the Composed Image Retrieval goal is to retrieve images visually similar to the reference one that integrates th…