8 papers
D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models
Bianca Raimondi, Davide Evangelista, Maurizio Gabbrielli +1
Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that appears to be internally represented by th…
From Reasoning to Code: GRPO Optimization for Underrepresented Languages
Federico Pennino, Bianca Raimondi, Massimo Rondelli +2
Generating accurate and executable code using Large Language Models (LLMs) remains a significant challenge for underrepresented programming languages, such as Prolog and Lisp, due…
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
Bianca Raimondi, Maurizio Gabbrielli
The black-box nature of Large Language Models necessitates novel evaluation frameworks that transcend surface-level performance metrics. This study investigates the internal neural…
The CompMath-MCQ Dataset: Are LLMs Ready for Higher-Level Math?
Bianca Raimondi, Francesco Pivi, Davide Evangelista +1
The evaluation of Large Language Models (LLMs) on mathematical reasoning has largely focused on elementary problems, competition-style questions, or formal theorem proving, leaving…
Learning Factors in AI-Augmented Education: A Comparative Study of Middle and High School Students
Gaia Ebli, Bianca Raimondi, Maurizio Gabbrielli
The increasing integration of AI tools in education has led prior research to explore their impact on learning processes. Nevertheless, most existing studies focus on higher educat…
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
Bianca Raimondi, Daniela Dalbagno, Maurizio Gabbrielli
Large language models (LLMs) have been shown to internalize human-like biases during finetuning, yet the mechanisms by which these biases manifest remain unclear. In this work, we…