4 papers
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
Konstantinos Voudouris, Mirko Thalmann, Alex Kipnis +2
Scientists, policy-makers, business leaders, and members of the public care about what modern artificial intelligence systems are disposed to do. Yet terms such as capabilities, pr…
metabeta -- A fast neural model for Bayesian mixed-effects regression
Alex Kipnis, Marcel Binz, Eric Schulz
Hierarchical data with multiple observations per group is ubiquitous in empirical sciences and is often analyzed using mixed-effects regression. In such models, Bayesian inference…
Centaur: a foundation model of human cognition
Marcel Binz, Elif Akata, Matthias Bethge +37
Establishing a unified theory of cognition has been a major goal of psychology. While there have been previous attempts to instantiate such theories by building computational model…
metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Alex Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff +1
Large Language Models (LLMs) vary in their abilities on a range of tasks. Initiatives such as the Open LLM Leaderboard aim to quantify these differences with several large benchmar…