most cited100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.AI2025

Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents

Irene Testini, José Hernández-Orallo, Lorenzo Pacchiardi

Data science aims to extract insights from data to support decision-making processes. Recently, Large Language Models (LLMs) have been increasingly used as assistants for data scie…

cs.AI2025

Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI

Danaja Rutar, Alva Markelius, Konstantinos Voudouris +2

One of the core components of our world models is 'intuitive physics' - an understanding of objects, space, and causality. This capability enables us to predict events, plan action…

cs.AI2025

General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed +23

Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…

cs.LG2025

What should an AI assessor optimise for?

Daniel Romero-Alvarado, Fernando Martínez-Plumed, José Hernández-Orallo

An AI assessor is an external, ideally indepen-dent system that predicts an indicator, e.g., a loss value, of another AI system. Assessors can lever-age information from the test r…

cs.CL20241 cited

100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances

Lorenzo Pacchiardi, Lucy G. Cheke, José Hernández-Orallo

Predicting the performance of LLMs on individual task instances is essential to ensure their reliability in high-stakes applications. To do so, a possibility is to evaluate the con…