1 citations · 1 across the 5 of their papers we have counts for
6 papers · 1 filter
Human agency in initial human-AI proof formalization workflows
Katherine M. Collins, Simon Frieder, Jonas Bayer +14
For centuries, human mathematicians have written proofs to substantiate their mathematical arguments; yet, the ability to automatically verify the validity of proofs has long been…
AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
Lance Ying, Ryan Truong, Prafull Sharma +9
Rigorously evaluating machine intelligence against the broad spectrum of human general intelligence has become increasingly important and challenging in this era of rapid technolog…
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
Irene Testini, José Hernández-Orallo, Lorenzo Pacchiardi
Data science aims to extract insights from data to support decision-making processes. Recently, Large Language Models (LLMs) have been increasingly used as assistants for data scie…
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
John Burden, Marko TeÅ¡iÄ, Lorenzo Pacchiardi +1
Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation pa…
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
Danaja Rutar, Alva Markelius, Konstantinos Voudouris +2
One of the core components of our world models is 'intuitive physics' - an understanding of objects, space, and causality. This capability enables us to predict events, plan action…
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Lexin Zhou, Lorenzo Pacchiardi, Fernando MartÃnez-Plumed +23
Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…