6 papers · 1 filter
Inferring Capabilities from Task Performance with Bayesian Triangulation
John Burden, Konstantinos Voudouris, Ryan Burnell +3
As machine learning models become more general, we need to characterise them in richer, more meaningful ways. We describe a method to infer the cognitive profile of a system from d…
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
John Burden, Marko TeÅ¡iÄ, Lorenzo Pacchiardi +1
Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation pa…
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Lexin Zhou, Lorenzo Pacchiardi, Fernando MartÃnez-Plumed +23
Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…
The Animal-AI Environment: A Virtual Laboratory For Comparative Cognition and Artificial Intelligence Research
Konstantinos Voudouris, Ibrahim Alhas, Wout Schellaert +11
The Animal-AI Environment is a unique game-based research platform designed to facilitate collaboration between the artificial intelligence and comparative cognition research commu…
Predictable Artificial Intelligence
Lexin Zhou, Pablo A. Moreno-Casares, Fernando MartÃnez-Plumed +12
We introduce the fundamental ideas and challenges of Predictable AI, a nascent research area that explores the ways in which we can anticipate key validity indicators (e.g., perfor…
Conversational Complexity for Assessing Risk in Large Language Models
John Burden, Manuel Cebrian, Jose Hernandez-Orallo
Large Language Models (LLMs) present a dual-use dilemma: they enable beneficial applications while harboring potential for harm, particularly through conversational interactions. D…