3 papers
Capabilities Ain't All You Need: Measuring Propensities in AI
Daniel Romero-Alvarado, Fernando Martínez-Plumed, Lorenzo Pacchiardi +11
AI evaluation has primarily focused on measuring capabilities, with formal approaches inspired from Item Response Theory (IRT) being increasingly applied. Yet propensities - the te…
From Human-Level AI Tales to AI Leveling Human Scales
Peter Romero, Fernando Martínez-Plumed, Zachary R. Tidler +11
Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose…
What should an AI assessor optimise for?
Daniel Romero-Alvarado, Fernando Martínez-Plumed, José Hernández-Orallo
An AI assessor is an external, ideally indepen-dent system that predicts an indicator, e.g., a loss value, of another AI system. Assessors can lever-age information from the test r…