6 papers
How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness
Polina Gordienko, Georg Schollmeyer, Frauke Kreuter +1
Multi-task benchmarks have become a central pillar of machine learning research, yet their growing influence has incentivised benchmark gaming -- strategic actions taken to improve…
Beyond Arrow: From Impossibility to Possibilities in Multi-Criteria Benchmarking
Polina Gordienko, Christoph Jansen, Julian Rodemann +1
Modern benchmarks such as HELM MMLU account for multiple metrics like accuracy, robustness and efficiency. When trying to turn these metrics into a single ranking, natural aggregat…
Empirical Decision Theory
Christoph Jansen, Georg Schollmeyer, Thomas Augustin +1
Analyzing decision problems under uncertainty commonly relies on idealizing assumptions about the describability of the world, with the most prominent examples being the closed wor…
Union-Free Generic Depth for Non-Standard Data
Hannah Blocher, Georg Schollmeyer
Non-standard data, which fall outside classical statistical data formats, challenge state-of-the-art analysis. Examples of non-standard data include partial orders and mixed catego…
Reciprocal Learning
Julian Rodemann, Christoph Jansen, Georg Schollmeyer
We demonstrate that a wide array of machine learning algorithms are specific instances of one single paradigm: reciprocal learning. These instances range from active learning over…
Data depth functions for non-standard data by use of formal concept analysis
Hannah Blocher, Georg Schollmeyer
In this article we introduce a notion of depth functions for data types that are not given in standard statistical data formats. We focus on data that cannot be represented by one…