5 papers
Different Facets of Verbalised Overconfidence: an Interpretability Study
Davide Mazzaccara, Leonardo Bertolazzi, Raffaella Bernardi
Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled reasoning scenarios that manipulate logica…
FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games
Leonardo Bertolazzi, Katya Tentori, Raffaella Bernardi
Large language models (LLMs) are increasingly deployed as autonomous agents in scientific tasks. Yet whether these systems can effectively engage in forms of inductive reasoning re…
Teaching Small Language Models to Learn Logic through Meta-Learning
Leonardo Bertolazzi, Manuel Vargas Guzmán, Raffaella Bernardi +2
Large language models (LLMs) are increasingly evaluated on reasoning tasks, yet their logical abilities remain contested. To address this, we study LLMs' reasoning in a well-define…
The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
Leonardo Bertolazzi, Philipp Mondorf, Barbara Plank +1
The ability of large language models (LLMs) to validate their output and identify potential errors is crucial for ensuring robustness and reliability. However, current research ind…
A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences
Leonardo Bertolazzi, Albert Gatt, Raffaella Bernardi
The reasoning abilities of Large Language Models (LLMs) are becoming a central focus of study in NLP. In this paper, we consider the case of syllogistic reasoning, an area of deduc…