activity
20222024
most citedMultimodal Emotion Recognition with Modality-Pairwise Unsupervised Contrastive Loss

3 citations · 6 across the 5 of their papers we have counts for

collaborators

5 papers

cs.AI2024

Automatic benchmarking of large multimodal models via iterative experiment programming

Alessandro Conti, Enrico Fini, Paolo Rota +3

Assessing the capabilities of large multimodal models (LMMs) often requires the creation of ad-hoc evaluations. Currently, building new benchmarks requires tremendous amounts of ma…

cs.CV20241 cited

Vocabulary-free Image Classification and Semantic Segmentation

Alessandro Conti, Enrico Fini, Massimiliano Mancini +3

Large vision-language models revolutionized image classification and semantic segmentation paradigms. However, they typically assume a pre-defined set of categories, or vocabulary,…

cs.CV2024

Test-Time Zero-Shot Temporal Action Localization

Benedetta Liberatori, Alessandro Conti, Paolo Rota +2

Zero-Shot Temporal Action Localization (ZS-TAL) seeks to identify and locate actions in untrimmed videos unseen during training. Existing ZS-TAL methods involve fine-tuning a model…

cs.CV20232 cited

The Unreasonable Effectiveness of Large Language-Vision Models for Source-free Video Domain Adaptation

Giacomo Zara, Alessandro Conti, Subhankar Roy +3

Source-Free Video Unsupervised Domain Adaptation (SFVUDA) task consists in adapting an action recognition model, trained on a labelled source dataset, to an unlabelled target datas…

cs.CV20223 cited

Multimodal Emotion Recognition with Modality-Pairwise Unsupervised Contrastive Loss

Riccardo Franceschini, Enrico Fini, Cigdem Beyan +3

Emotion recognition is involved in several real-world applications. With an increase in available modalities, automatic understanding of emotions is being performed more accurately…