4 papers
[De|Re]constructing VLMs' Reasoning in Counting
Simone Alghisi, Gabriel Roccabruna, Massimo Rizzoli +2
Vision-Language Models (VLMs) have recently gained attention due to their competitive performance on multiple downstream tasks, achieved by following user-input instructions. Howev…
CIVET: Systematic Evaluation of Understanding in VLMs
Massimo Rizzoli, Simone Alghisi, Olha Khomyn +3
While Vision-Language Models (VLMs) have achieved competitive performance in various tasks, their comprehension of the underlying structure and semantics of a scene remains underst…
Will LLMs Replace the Encoder-Only Models in Temporal Relation Classification?
Gabriel Roccabruna, Massimo Rizzoli, Giuseppe Riccardi
The automatic detection of temporal relations among events has been mainly investigated with encoder-only models such as RoBERTa. Large Language Models (LLM) have recently shown pr…
Are LLMs Robust for Spoken Dialogues?
Seyed Mahed Mousavi, Gabriel Roccabruna, Simone Alghisi +3
Large Pre-Trained Language Models have demonstrated state-of-the-art performance in different downstream tasks, including dialogue state tracking and end-to-end response generation…