1 citations · 2 across the 2 of their papers we have counts for
3 papers
I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs
John Burden, Jonathan Prunty, Ben Slater +3
Multimodal large language models (MLLMs) achieve strong performance on vision-language tasks, yet their visual processing is opaque. Most black-box evaluations measure task accurac…
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
Lorenzo Pacchiardi, Lucy G. Cheke, José Hernández-Orallo
Predicting the performance of LLMs on individual task instances is essential to ensure their reliability in high-stakes applications. To do so, a possibility is to evaluate the con…
The Animal-AI Environment: Training and Testing Animal-Like Artificial Cognition
Benjamin Beyret, José Hernández-Orallo, Lucy Cheke +3
Recent advances in artificial intelligence have been strongly driven by the use of game environments for training and evaluating agents. Games are often accessible and versatile, w…