activity
20212023
most citedMultimodal Across Domains Gaze Target Detection

25 citations · 37 across the 13 of their papers we have counts for

collaborators

29 papers

cs.CV2024

Retrieval-enriched zero-shot image classification in low-resource domains

Nicola Dall'Asen, Yiming Wang, Enrico Fini +1

Low-resource domains, characterized by scarce data and annotations, present significant challenges for language and visual understanding tasks, with the latter much under-explored…

cs.CV20241 cited

The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models

Simone Caldarella, Massimiliano Mancini, Elisa Ricci +1

Vision-Language Models (VLMs) combine visual and textual understanding, rendering them well-suited for diverse tasks like generating image captions and answering visual questions a…

cs.CV2024

Conditioned Prompt-Optimization for Continual Deepfake Detection

Francesco Laiti, Benedetta Liberatori, Thomas De Min +1

The rapid advancement of generative models has significantly enhanced the realism and customization of digital content creation. The increasing power of these tools, coupled with t…

cs.AI2024

Automatic benchmarking of large multimodal models via iterative experiment programming

Alessandro Conti, Enrico Fini, Paolo Rota +3

Assessing the capabilities of large multimodal models (LMMs) often requires the creation of ad-hoc evaluations. Currently, building new benchmarks requires tremendous amounts of ma…

cs.CV20241 cited

Vocabulary-free Image Classification and Semantic Segmentation

Alessandro Conti, Enrico Fini, Massimiliano Mancini +3

Large vision-language models revolutionized image classification and semantic segmentation paradigms. However, they typically assume a pre-defined set of categories, or vocabulary,…

cs.CV2024

Test-Time Zero-Shot Temporal Action Localization

Benedetta Liberatori, Alessandro Conti, Paolo Rota +2

Zero-Shot Temporal Action Localization (ZS-TAL) seeks to identify and locate actions in untrimmed videos unseen during training. Existing ZS-TAL methods involve fine-tuning a model…