activity
20242026
collaborators

6 papers

cs.LG2026

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness

Polina Gordienko, Georg Schollmeyer, Frauke Kreuter +1

Multi-task benchmarks have become a central pillar of machine learning research, yet their growing influence has incentivised benchmark gaming -- strategic actions taken to improve…

cs.LG2026

Beyond Arrow: From Impossibility to Possibilities in Multi-Criteria Benchmarking

Polina Gordienko, Christoph Jansen, Julian Rodemann +1

Modern benchmarks such as HELM MMLU account for multiple metrics like accuracy, robustness and efficiency. When trying to turn these metrics into a single ranking, natural aggregat…

stat.ME2025

Empirical Decision Theory

Christoph Jansen, Georg Schollmeyer, Thomas Augustin +1

Analyzing decision problems under uncertainty commonly relies on idealizing assumptions about the describability of the world, with the most prominent examples being the closed wor…

stat.ME2024

Union-Free Generic Depth for Non-Standard Data

Hannah Blocher, Georg Schollmeyer

Non-standard data, which fall outside classical statistical data formats, challenge state-of-the-art analysis. Examples of non-standard data include partial orders and mixed catego…

stat.ML2024

Reciprocal Learning

Julian Rodemann, Christoph Jansen, Georg Schollmeyer

We demonstrate that a wide array of machine learning algorithms are specific instances of one single paradigm: reciprocal learning. These instances range from active learning over…

math.ST2024

Data depth functions for non-standard data by use of formal concept analysis

Hannah Blocher, Georg Schollmeyer

In this article we introduce a notion of depth functions for data types that are not given in standard statistical data formats. We focus on data that cannot be represented by one…