5 papers
How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness
Polina Gordienko, Georg Schollmeyer, Frauke Kreuter +1
Multi-task benchmarks have become a central pillar of machine learning research, yet their growing influence has incentivised benchmark gaming -- strategic actions taken to improve…
Beyond Arrow: From Impossibility to Possibilities in Multi-Criteria Benchmarking
Polina Gordienko, Christoph Jansen, Julian Rodemann +1
Modern benchmarks such as HELM MMLU account for multiple metrics like accuracy, robustness and efficiency. When trying to turn these metrics into a single ranking, natural aggregat…
Union-Free Generic Depth for Non-Standard Data
Hannah Blocher, Georg Schollmeyer
Non-standard data, which fall outside classical statistical data formats, challenge state-of-the-art analysis. Examples of non-standard data include partial orders and mixed catego…
Reciprocal Learning
Julian Rodemann, Christoph Jansen, Georg Schollmeyer
We demonstrate that a wide array of machine learning algorithms are specific instances of one single paradigm: reciprocal learning. These instances range from active learning over…
Statistical Multicriteria Benchmarking via the GSD-Front
Christoph Jansen, Georg Schollmeyer, Julian Rodemann +2
Given the vast number of classifiers that have been (and continue to be) proposed, reliable methods for comparing them are becoming increasingly important. The desire for reliabili…