3 papers
cs.CL2025
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It
Seyed Mahed Mousavi, Edoardo Cecchinato, Lucia Hornikova +1
We conduct a systematic audit of three widely used reasoning benchmarks, SocialIQa, FauxPas-EAI, and ToMi, and uncover pervasive flaws in both benchmark items and evaluation method…
cs.CV2025
CIVET: Systematic Evaluation of Understanding in VLMs
Massimo Rizzoli, Simone Alghisi, Olha Khomyn +3
While Vision-Language Models (VLMs) have achieved competitive performance in various tasks, their comprehension of the underlying structure and semantics of a scene remains underst…
cs.CL2024
Are LLMs Robust for Spoken Dialogues?
Seyed Mahed Mousavi, Gabriel Roccabruna, Simone Alghisi +3
Large Pre-Trained Language Models have demonstrated state-of-the-art performance in different downstream tasks, including dialogue state tracking and end-to-end response generation…