8 papers
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
Ben Slater, Matteo G. Mecattaf, Lucy G. Cheke +2
Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in increasingly agentic and autono…
The Missing Red Line: How Commercial Pressure Erodes AI Safety Boundaries
Nora Petrova, John Burden
What happens when an AI assistant is told to "maximise sales" while a user asks about drug interactions? We find that commercial system prompts can override safety training, causin…
Inferring Capabilities from Task Performance with Bayesian Triangulation
John Burden, Konstantinos Voudouris, Ryan Burnell +3
As machine learning models become more general, we need to characterise them in richer, more meaningful ways. We describe a method to infer the cognitive profile of a system from d…
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
John Burden, Marko TeÅ¡iÄ, Lorenzo Pacchiardi +1
Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation pa…
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Lexin Zhou, Lorenzo Pacchiardi, Fernando MartÃnez-Plumed +23
Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…
The Animal-AI Environment: A Virtual Laboratory For Comparative Cognition and Artificial Intelligence Research
Konstantinos Voudouris, Ibrahim Alhas, Wout Schellaert +11
The Animal-AI Environment is a unique game-based research platform designed to facilitate collaboration between the artificial intelligence and comparative cognition research commu…