25 citations · 37 across the 13 of their papers we have counts for
29 papers
Retrieval-enriched zero-shot image classification in low-resource domains
Nicola Dall'Asen, Yiming Wang, Enrico Fini +1
Low-resource domains, characterized by scarce data and annotations, present significant challenges for language and visual understanding tasks, with the latter much under-explored…
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
Simone Caldarella, Massimiliano Mancini, Elisa Ricci +1
Vision-Language Models (VLMs) combine visual and textual understanding, rendering them well-suited for diverse tasks like generating image captions and answering visual questions a…
Conditioned Prompt-Optimization for Continual Deepfake Detection
Francesco Laiti, Benedetta Liberatori, Thomas De Min +1
The rapid advancement of generative models has significantly enhanced the realism and customization of digital content creation. The increasing power of these tools, coupled with t…
Automatic benchmarking of large multimodal models via iterative experiment programming
Alessandro Conti, Enrico Fini, Paolo Rota +3
Assessing the capabilities of large multimodal models (LMMs) often requires the creation of ad-hoc evaluations. Currently, building new benchmarks requires tremendous amounts of ma…
Vocabulary-free Image Classification and Semantic Segmentation
Alessandro Conti, Enrico Fini, Massimiliano Mancini +3
Large vision-language models revolutionized image classification and semantic segmentation paradigms. However, they typically assume a pre-defined set of categories, or vocabulary,…
Test-Time Zero-Shot Temporal Action Localization
Benedetta Liberatori, Alessandro Conti, Paolo Rota +2
Zero-Shot Temporal Action Localization (ZS-TAL) seeks to identify and locate actions in untrimmed videos unseen during training. Existing ZS-TAL methods involve fine-tuning a model…