activity
20242026
collaborators

5 papers

cs.CL2026

Corpus Prevalence of Multiple-Choice Question Options

Leonidas Zotos, Hedderik van Rijn, Malvina Nissim

In recent years, corpus-driven AI methods, such as Large Language Models (LLMs), have seen widespread use in education. While on the surface their abilities look promising for task…

cs.LG2025

EAGer: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling

Daniel Scalena, Leonidas Zotos, Elisabetta Fersini +2

With the rise of reasoning language models and test-time scaling methods as a paradigm for improving model performance, substantial computation is often required to generate multip…

cs.CL2025

NLP Methods May Actually Be Better Than Professors at Estimating Question Difficulty

Leonidas Zotos, Ivo Pascal de Jong, Matias Valdenegro-Toro +3

Estimating the difficulty of exam questions is essential for developing good exams, but professors are not always good at this task. We compare various Large Language Model-based m…

cs.CL2024

Are You Doubtful? Oh, It Might Be Difficult Then! Exploring the Use of Model Uncertainty for Question Difficulty Estimation

Leonidas Zotos, Hedderik van Rijn, Malvina Nissim

In an educational setting, an estimate of the difficulty of multiple-choice questions (MCQs), a commonly used strategy to assess learning progress, constitutes very useful informat…

cs.CL2024

Can Model Uncertainty Function as a Proxy for Multiple-Choice Question Item Difficulty?

Leonidas Zotos, Hedderik van Rijn, Malvina Nissim

Estimating the difficulty of multiple-choice questions would be great help for educators who must spend substantial time creating and piloting stimuli for their tests, and for lear…