activity
20242026
collaborators

7 papers

cs.LG2026

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness

Polina Gordienko, Georg Schollmeyer, Frauke Kreuter +1

Multi-task benchmarks have become a central pillar of machine learning research, yet their growing influence has incentivised benchmark gaming -- strategic actions taken to improve…

cs.LG2026

Beyond Arrow: From Impossibility to Possibilities in Multi-Criteria Benchmarking

Polina Gordienko, Christoph Jansen, Julian Rodemann +1

Modern benchmarks such as HELM MMLU account for multiple metrics like accuracy, robustness and efficiency. When trying to turn these metrics into a single ranking, natural aggregat…

cs.CL2025

Statistical Multicriteria Evaluation of LLM-Generated Text

Esteban Garces Arias, Hannah Blocher, Julian Rodemann +2

Assessing the quality of LLM-generated text remains a fundamental challenge in natural language processing. Current evaluation approaches often rely on isolated metrics or simplist…

cs.AI2025

A Statistical Case Against Empirical Human-AI Alignment

Julian Rodemann, Esteban Garces Arias, Christoph Luther +2

Empirical human-AI alignment aims to make AI systems act in line with observed human behavior. While noble in its goals, we argue that empirical alignment can inadvertently introdu…

stat.ML2024

Reciprocal Learning

Julian Rodemann, Christoph Jansen, Georg Schollmeyer

We demonstrate that a wide array of machine learning algorithms are specific instances of one single paradigm: reciprocal learning. These instances range from active learning over…

stat.ML2024

Statistical Multicriteria Benchmarking via the GSD-Front

Christoph Jansen, Georg Schollmeyer, Julian Rodemann +2

Given the vast number of classifiers that have been (and continue to be) proposed, reliable methods for comparing them are becoming increasingly important. The desire for reliabili…