6 papers
Exploring Geographic Relative Space in Large Language Models through Activation Patching
Stef De Sabbata, Rahul Baiju, Stefano Mizzaro +1
The increased use of Large Language Models (LLMs) in geography raises substantial questions about the safety of integrating these tools across a wide range of processes and analyse…
Beyond Seeing Is Believing: On Crowdsourced Detection of Audiovisual Deepfakes
Michael Soprano, Andrea Cioci, Stefano Mizzaro
Deepfakes are increasingly realistic and easy to produce, raising concerns about the reliability of human judgments in misinformation settings. We study audiovisual deepfake detect…
The Effect of Document Summarization on LLM-Based Relevance Judgments
Samaneh Mohtadi, Kevin Roitero, Stefano Mizzaro +1
Relevance judgments are central to the evaluation of Information Retrieval (IR) systems, but obtaining them from human annotators is costly and time-consuming. Large Language Model…
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
Riccardo Lunardi, Vincenzo Della Mea, Stefano Mizzaro +1
Large Language Models (LLMs) effectiveness is usually evaluated by means of benchmarks such as MMLU, ARC-C, or HellaSwag, where questions are presented in their original wording, t…
Geospatial Mechanistic Interpretability of Large Language Models
Stef De Sabbata, Stefano Mizzaro, Kevin Roitero
Large Language Models (LLMs) have demonstrated unprecedented capabilities across various natural language processing tasks. Their ability to process and generate viable text and co…
Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
Kevin Roitero, Dustin Wright, Michael Soprano +2
Evaluating the truthfulness of online content is critical for combating misinformation. This study examines the efficiency and effectiveness of crowdsourced truthfulness assessment…