4 papers
Visuospatial Perspective Taking in Multimodal Language Models
Jonathan Prunty, Seraphina Zhang, Patrick Quinn +3
As multimodal language models (MLMs) are increasingly used in social and collaborative settings, it is crucial to evaluate their perspective-taking abilities. Existing benchmarks l…
Capabilities Ain't All You Need: Measuring Propensities in AI
Daniel Romero-Alvarado, Fernando MartÃnez-Plumed, Lorenzo Pacchiardi +11
AI evaluation has primarily focused on measuring capabilities, with formal approaches inspired from Item Response Theory (IRT) being increasingly applied. Yet propensities - the te…
I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
John Burden, Jonathan Prunty, Ben Slater +3
Multimodal large language models (MLLMs) achieve strong performance on vision-language tasks, yet their visual processing is opaque. Most black-box evaluations measure task accurac…
A little less conversation, a little more action, please: Investigating the physical common-sense of LLMs in a 3D embodied environment
Matteo G. Mecattaf, Ben Slater, Marko TeÅ¡iÄ +3
As general-purpose tools, Large Language Models (LLMs) must often reason about everyday physical environments. In a question-and-answer capacity, understanding the interactions of…