1 citations · 1 across the 2 of their papers we have counts for
5 papers
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
Irene Testini, José Hernández-Orallo, Lorenzo Pacchiardi
Data science aims to extract insights from data to support decision-making processes. Recently, Large Language Models (LLMs) have been increasingly used as assistants for data scie…
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
Danaja Rutar, Alva Markelius, Konstantinos Voudouris +2
One of the core components of our world models is 'intuitive physics' - an understanding of objects, space, and causality. This capability enables us to predict events, plan action…
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed +23
Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…
What should an AI assessor optimise for?
Daniel Romero-Alvarado, Fernando Martínez-Plumed, José Hernández-Orallo
An AI assessor is an external, ideally indepen-dent system that predicts an indicator, e.g., a loss value, of another AI system. Assessors can lever-age information from the test r…
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
Lorenzo Pacchiardi, Lucy G. Cheke, José Hernández-Orallo
Predicting the performance of LLMs on individual task instances is essential to ensure their reliability in high-stakes applications. To do so, a possibility is to evaluate the con…