3 papers
cs.CL2025
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
Riccardo Lunardi, Vincenzo Della Mea, Stefano Mizzaro +1
Large Language Models (LLMs) effectiveness is usually evaluated by means of benchmarks such as MMLU, ARC-C, or HellaSwag, where questions are presented in their original wording, t…
cs.CL2025
Political Ideology Shifts in Large Language Models
Pietro Bernardelle, Stefano Civelli, Leon Fröhling +3
Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideological biases. In this work, we examin…
cs.CL2024
Mapping and Influencing the Political Ideology of Large Language Models using Synthetic Personas
Pietro Bernardelle, Leon Fröhling, Stefano Civelli +3
The analysis of political biases in large language models (LLMs) has primarily examined these systems as single entities with fixed viewpoints. While various methods exist for meas…