5 papers
Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs
Andrea Bacciu, Andrea Alfarano, Saab Mansour +2
Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale e…
Select, Label, Evaluate: Active Testing in NLP
Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu +2
Human annotation cost and time remain significant bottlenecks in Natural Language Processing (NLP), with test data annotation being particularly expensive due to the stringent requ…
Cross-Lingual LLM-Judge Transfer via Evaluation Decomposition
Ivaxi Sheth, Zeno Jonke, Amin Mantrach +1
As large language models are increasingly deployed across diverse real-world applications, extending automated evaluation beyond English has become a critical challenge. Existing e…
Multilingual Self-Taught Faithfulness Evaluators
Carlo Alfano, Aymen Al Marjani, Zeno Jonke +3
The growing use of large language models (LLMs) has increased the need for automatic evaluation systems, particularly to address the challenge of information hallucination. Althoug…
Monte Carlo Temperature: a robust sampling strategy for LLM's uncertainty quantification methods
Nicola Cecere, Andrea Bacciu, Ignacio Fernández TobÃas +1
Uncertainty quantification (UQ) in Large Language Models (LLMs) is essential for their safe and reliable deployment, particularly in critical applications where incorrect outputs c…