85 citations · 107 across the 11 of their papers we have counts for
13 papers
On Large Multimodal Models as Open-World Image Classifiers
Alessandro Conti, Massimiliano Mancini, Enrico Fini +3
Traditional image classification requires a predefined list of semantic categories. In contrast, Large Multimodal Models (LMMs) can sidestep this requirement by classifying images…
Retrieval-enriched zero-shot image classification in low-resource domains
Nicola Dall'Asen, Yiming Wang, Enrico Fini +1
Low-resource domains, characterized by scarce data and annotations, present significant challenges for language and visual understanding tasks, with the latter much under-explored…
Automatic benchmarking of large multimodal models via iterative experiment programming
Alessandro Conti, Enrico Fini, Paolo Rota +3
Assessing the capabilities of large multimodal models (LMMs) often requires the creation of ad-hoc evaluations. Currently, building new benchmarks requires tremendous amounts of ma…
Vocabulary-free Image Classification and Semantic Segmentation
Alessandro Conti, Enrico Fini, Massimiliano Mancini +3
Large vision-language models revolutionized image classification and semantic segmentation paradigms. However, they typically assume a pre-defined set of categories, or vocabulary,…
Semi-supervised learning made simple with self-supervised clustering
Enrico Fini, Pietro Astolfi, Karteek Alahari +4
Self-supervised learning models have been shown to learn rich visual representations without requiring human annotations. However, in many real-world scenarios, labels are partiall…
Vocabulary-free Image Classification
Alessandro Conti, Enrico Fini, Massimiliano Mancini +3
Recent advances in large vision-language models have revolutionized the image classification paradigm. Despite showing impressive zero-shot capabilities, a pre-defined set of categ…