11 papers
Human agency in initial human-AI proof formalization workflows
Katherine M. Collins, Simon Frieder, Jonas Bayer +14
For centuries, human mathematicians have written proofs to substantiate their mathematical arguments; yet, the ability to automatically verify the validity of proofs has long been…
Multi-agent AI systems outperform human teams in creativity
Tiancheng Hu, Yixuan Jiang, Haotian Li +5
Although artificial intelligence (AI) now matches or exceeds human performance across numerous cognitive tasks, creativity remains a highly contested frontier. As AI systems based…
From Human-Level AI Tales to AI Leveling Human Scales
Peter Romero, Fernando MartÃnez-Plumed, Zachary R. Tidler +11
Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose…
AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
Lance Ying, Ryan Truong, Prafull Sharma +9
Rigorously evaluating machine intelligence against the broad spectrum of human general intelligence has become increasingly important and challenging in this era of rapid technolog…
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
Konstantinos Voudouris, Mirko Thalmann, Alex Kipnis +2
Scientists, policy-makers, business leaders, and members of the public care about what modern artificial intelligence systems are disposed to do. Yet terms such as capabilities, pr…
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
Irene Testini, José Hernández-Orallo, Lorenzo Pacchiardi
Data science aims to extract insights from data to support decision-making processes. Recently, Large Language Models (LLMs) have been increasingly used as assistants for data scie…