10 papers
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
Ben Slater, Matteo G. Mecattaf, Lucy G. Cheke +2
Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in increasingly agentic and autono…
Visuospatial Perspective Taking in Multimodal Language Models
Jonathan Prunty, Seraphina Zhang, Patrick Quinn +3
As multimodal language models (MLMs) are increasingly used in social and collaborative settings, it is crucial to evaluate their perspective-taking abilities. Existing benchmarks l…
I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
John Burden, Jonathan Prunty, Ben Slater +3
Multimodal large language models (MLLMs) achieve strong performance on vision-language tasks, yet their visual processing is opaque. Most black-box evaluations measure task accurac…
Inferring Capabilities from Task Performance with Bayesian Triangulation
John Burden, Konstantinos Voudouris, Ryan Burnell +3
As machine learning models become more general, we need to characterise them in richer, more meaningful ways. We describe a method to infer the cognitive profile of a system from d…
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
Danaja Rutar, Alva Markelius, Konstantinos Voudouris +2
One of the core components of our world models is 'intuitive physics' - an understanding of objects, space, and causality. This capability enables us to predict events, plan action…
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Lexin Zhou, Lorenzo Pacchiardi, Fernando MartÃnez-Plumed +23
Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…