most cited100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances

1 citations · 1 across the 1 of their papers we have counts for

collaborators

5 papers

cs.AI2025

Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents

Irene Testini, José Hernández-Orallo, Lorenzo Pacchiardi

Data science aims to extract insights from data to support decision-making processes. Recently, Large Language Models (LLMs) have been increasingly used as assistants for data scie…

cs.AI2025

General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed +23

Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…

cs.AI2025

Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture

John Burden, Marko Tešić, Lorenzo Pacchiardi +1

Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation pa…

cs.CL2025

PredictaBoard: Benchmarking LLM Score Predictability

Lorenzo Pacchiardi, Konstantinos Voudouris, Ben Slater +4

Despite possessing impressive skills, Large Language Models (LLMs) often fail unpredictably, demonstrating inconsistent success in even basic common sense reasoning tasks. This unp…

cs.CL20241 cited

100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances

Lorenzo Pacchiardi, Lucy G. Cheke, José Hernández-Orallo

Predicting the performance of LLMs on individual task instances is essential to ensure their reliability in high-stakes applications. To do so, a possibility is to evaluate the con…