5 papers
SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models
Sebastiano Monti, Carlo Nicolini, Gianni Pellegrini +2
Although the capabilities of large language models have been increasingly tested on complex reasoning tasks, their long-horizon planning abilities have not yet been extensively inv…
ProfVLM: A lightweight video-language model for multi-view proficiency estimation
Edoardo Bianchi, Jacopo Staiano, Antonio Liotta
Most existing approaches formulate action quality assessment and skill proficiency estimation as discriminative prediction tasks, typically producing discrete labels or scores with…
MedSyn: Enhancing Diagnostics with Human-AI Collaboration
Burcu Sayin, Ipek Baris Schlicht, Ngoc Vo Hong +4
Clinical decision-making is inherently complex, often influenced by cognitive biases, incomplete information, and case ambiguity. Large Language Models (LLMs) have shown promise as…
The LLM Wears Prada: Analysing Gender Bias and Stereotypes through Online Shopping Data
Massimiliano Luca, Ciro Beneduce, Bruno Lepri +1
With the wide and cross-domain adoption of Large Language Models, it becomes crucial to assess to which extent the statistical correlations in training data, which underlie their i…
Face the Facts! Evaluating RAG-based Pipelines for Professional Fact-Checking
Daniel Russo, Stefano Menini, Jacopo Staiano +1
Natural Language Processing and Generation systems have recently shown the potential to complement and streamline the costly and time-consuming job of professional fact-checkers. I…