activity
20242026
collaborators

5 papers

cs.AI2026

SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models

Sebastiano Monti, Carlo Nicolini, Gianni Pellegrini +2

Although the capabilities of large language models have been increasingly tested on complex reasoning tasks, their long-horizon planning abilities have not yet been extensively inv…

cs.CV2025

ProfVLM: A lightweight video-language model for multi-view proficiency estimation

Edoardo Bianchi, Jacopo Staiano, Antonio Liotta

Most existing approaches formulate action quality assessment and skill proficiency estimation as discriminative prediction tasks, typically producing discrete labels or scores with…

cs.LG2025

MedSyn: Enhancing Diagnostics with Human-AI Collaboration

Burcu Sayin, Ipek Baris Schlicht, Ngoc Vo Hong +4

Clinical decision-making is inherently complex, often influenced by cognitive biases, incomplete information, and case ambiguity. Large Language Models (LLMs) have shown promise as…

cs.AI2025

The LLM Wears Prada: Analysing Gender Bias and Stereotypes through Online Shopping Data

Massimiliano Luca, Ciro Beneduce, Bruno Lepri +1

With the wide and cross-domain adoption of Large Language Models, it becomes crucial to assess to which extent the statistical correlations in training data, which underlie their i…

cs.CL2024

Face the Facts! Evaluating RAG-based Pipelines for Professional Fact-Checking

Daniel Russo, Stefano Menini, Jacopo Staiano +1

Natural Language Processing and Generation systems have recently shown the potential to complement and streamline the costly and time-consuming job of professional fact-checkers. I…