4 papers
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
Tom Lucas, Alessio Buscemi, Alfredo Capozucca +2
Assessing whether Large Language Models outputs are factually grounded, epistemically calibrated, and methodologically reproducible is a prerequisite for responsible AI deployment.…
The Sandbox Configurator: A Framework to Support Technical Assessment in AI Regulatory Sandboxes
Alessio Buscemi, Thibault Simonetto, Daniele Pagani +3
The systematic assessment of AI systems is increasingly vital as these technologies enter high-stakes domains. To address this, the EU's Artificial Intelligence Act introduces AI R…
Evaluating Open-Source Large Language Models for Technical Telecom Question Answering
Arina Caraus, Alessio Buscemi, Sumit Kumar +1
Large Language Models (LLMs) have shown remarkable capabilities across various fields. However, their performance in technical domains such as telecommunications remains underexplo…
Mind the Language Gap: Automated and Augmented Evaluation of Bias in LLMs for High- and Low-Resource Languages
Alessio Buscemi, Cédric Lothritz, Sergio Morales +4
Large Language Models (LLMs) have exhibited impressive natural language processing capabilities but often perpetuate social biases inherent in their training data. To address this,…