most citedExploiting CLIP-based Multi-modal Approach for Artwork Classification and Retrieval

7 citations · 8 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

Improving Zero-shot Generalization of Learned Prompts via Unsupervised Knowledge Distillation

Marco Mistretta, Alberto Baldrati, Marco Bertini +1

Vision-Language Models (VLMs) demonstrate remarkable zero-shot generalization to unseen tasks, but fall short of the performance of supervised methods in generalizing to downstream…

cs.CV2024

Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing

Alberto Baldrati, Davide Morelli, Marcella Cornia +2

Fashion illustration is a crucial medium for designers to convey their creative vision and transform design concepts into tangible representations that showcase the interplay betwe…

cs.CV2023

Mapping Memes to Words for Multimodal Hateful Meme Classification

Giovanni Burbi, Alberto Baldrati, Lorenzo Agnolucci +2

Multimodal image-text memes are prevalent on the internet, serving as a unique form of communication that combines visual and textual elements to convey humor, ideas, or emotions.…

cs.CV20237 cited

Exploiting CLIP-based Multi-modal Approach for Artwork Classification and Retrieval

Alberto Baldrati, Marco Bertini, Tiberio Uricchio +1

Given the recent advances in multimodal image pretraining where visual models trained with semantically dense textual supervision tend to have better generalization capabilities th…

cs.CV2023

OpenFashionCLIP: Vision-and-Language Contrastive Learning with Open-Source Fashion Data

Giuseppe Cartella, Alberto Baldrati, Davide Morelli +3

The inexorable growth of online shopping and e-commerce demands scalable and robust machine learning-based solutions to accommodate customer requirements. In the context of automat…

cs.CV20231 cited

Composed Image Retrieval using Contrastive Learning and Task-oriented CLIP-based Features

Alberto Baldrati, Marco Bertini, Tiberio Uricchio +1

Given a query composed of a reference image and a relative caption, the Composed Image Retrieval goal is to retrieve images visually similar to the reference one that integrates th…