282 citations · 327 across the 52 of their papers we have counts for
7 papers · 1 filter
Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models
Samuele Poppi, Tobia Poppi, Federico Cocchi +3
Large-scale vision-and-language models, such as CLIP, are typically trained on web-scale data, which can introduce inappropriate content and lead to the development of unsafe and b…
With a Little Help from your own Past: Prototypical Memory Networks for Image Captioning
Manuele Barraco, Sara Sarto, Marcella Cornia +2
Image captioning, like many tasks involving vision and language, currently relies on Transformer-based architectures for extracting the semantics in an image and translating it int…
Let's ViCE! Mimicking Human Cognitive Behavior in Image Generation Evaluation
Federico Betti, Jacopo Staiano, Lorenzo Baraldi +2
Research in Image Generation has recently made significant progress, particularly boosted by the introduction of Vision-Language models which are able to produce high-quality visua…
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
Lorenzo Baraldi, Roberto Amoroso, Marcella Cornia +2
The use of self-supervised pre-training has emerged as a promising approach to enhance the performance of many different visual tasks. In this context, recent approaches have emplo…
Evaluating Synthetic Pre-Training for Handwriting Processing Tasks
Vittorio Pippi, Silvia Cascianelli, Lorenzo Baraldi +1
In this work, we explore massive pre-training on synthetic word images for enhancing the performance on four benchmark downstream handwriting analysis tasks. To this end, we build…
Multi-Class Unlearning for Image Classification via Weight Filtering
Samuele Poppi, Sara Sarto, Marcella Cornia +2
Machine Unlearning is an emerging paradigm for selectively removing the impact of training datapoints from a network. Unlike existing methods that target a limited subset or a sing…