5 papers
Capabilities Ain't All You Need: Measuring Propensities in AI
Daniel Romero-Alvarado, Fernando Martínez-Plumed, Lorenzo Pacchiardi +11
AI evaluation has primarily focused on measuring capabilities, with formal approaches inspired from Item Response Theory (IRT) being increasingly applied. Yet propensities - the te…
From Human-Level AI Tales to AI Leveling Human Scales
Peter Romero, Fernando Martínez-Plumed, Zachary R. Tidler +11
Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose…
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed +23
Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…
PredictaBoard: Benchmarking LLM Score Predictability
Lorenzo Pacchiardi, Konstantinos Voudouris, Ben Slater +4
Despite possessing impressive skills, Large Language Models (LLMs) often fail unpredictably, demonstrating inconsistent success in even basic common sense reasoning tasks. This unp…
What should an AI assessor optimise for?
Daniel Romero-Alvarado, Fernando Martínez-Plumed, José Hernández-Orallo
An AI assessor is an external, ideally indepen-dent system that predicts an indicator, e.g., a loss value, of another AI system. Assessors can lever-age information from the test r…