7 papers
How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness
Polina Gordienko, Georg Schollmeyer, Frauke Kreuter +1
Multi-task benchmarks have become a central pillar of machine learning research, yet their growing influence has incentivised benchmark gaming -- strategic actions taken to improve…
Beyond Arrow: From Impossibility to Possibilities in Multi-Criteria Benchmarking
Polina Gordienko, Christoph Jansen, Julian Rodemann +1
Modern benchmarks such as HELM MMLU account for multiple metrics like accuracy, robustness and efficiency. When trying to turn these metrics into a single ranking, natural aggregat…
Statistical Multicriteria Evaluation of LLM-Generated Text
Esteban Garces Arias, Hannah Blocher, Julian Rodemann +2
Assessing the quality of LLM-generated text remains a fundamental challenge in natural language processing. Current evaluation approaches often rely on isolated metrics or simplist…
A Statistical Case Against Empirical Human-AI Alignment
Julian Rodemann, Esteban Garces Arias, Christoph Luther +2
Empirical human-AI alignment aims to make AI systems act in line with observed human behavior. While noble in its goals, we argue that empirical alignment can inadvertently introdu…
Reciprocal Learning
Julian Rodemann, Christoph Jansen, Georg Schollmeyer
We demonstrate that a wide array of machine learning algorithms are specific instances of one single paradigm: reciprocal learning. These instances range from active learning over…
Statistical Multicriteria Benchmarking via the GSD-Front
Christoph Jansen, Georg Schollmeyer, Julian Rodemann +2
Given the vast number of classifiers that have been (and continue to be) proposed, reliable methods for comparing them are becoming increasingly important. The desire for reliabili…