activity
20242026
collaborators
Showing 2025Show all

5 papers · 1 filter

cs.AI2025

Inferring Capabilities from Task Performance with Bayesian Triangulation

John Burden, Konstantinos Voudouris, Ryan Burnell +3

As machine learning models become more general, we need to characterise them in richer, more meaningful ways. We describe a method to infer the cognitive profile of a system from d…

cs.AI2025

Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture

John Burden, Marko Tešić, Lorenzo Pacchiardi +1

Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation pa…

cs.AI2025

General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed +23

Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…

cs.AI2025

The Animal-AI Environment: A Virtual Laboratory For Comparative Cognition and Artificial Intelligence Research

Konstantinos Voudouris, Ibrahim Alhas, Wout Schellaert +11

The Animal-AI Environment is a unique game-based research platform designed to facilitate collaboration between the artificial intelligence and comparative cognition research commu…

cs.AI2025

Predictable Artificial Intelligence

Lexin Zhou, Pablo A. Moreno-Casares, Fernando Martínez-Plumed +12

We introduce the fundamental ideas and challenges of Predictable AI, a nascent research area that explores the ways in which we can anticipate key validity indicators (e.g., perfor…