Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness
Polina Gordienko, Georg Schollmeyer, Frauke Kreuter +1
Multi-task benchmarks have become a central pillar of machine learning research, yet their growing influence has incentivised benchmark gaming -- strategic actions taken to improve…
cs.LG2026
Beyond Arrow: From Impossibility to Possibilities in Multi-Criteria Benchmarking
Polina Gordienko, Christoph Jansen, Julian Rodemann +1
Modern benchmarks such as HELM MMLU account for multiple metrics like accuracy, robustness and efficiency. When trying to turn these metrics into a single ranking, natural aggregat…
cs.LG2023
Comparing Machine Learning Algorithms by Union-Free Generic Depth
Hannah Blocher, Georg Schollmeyer, Malte Nalenz +1
We propose a framework for descriptively analyzing sets of partial orders based on the concept of depth functions. Despite intensive studies in linear and metric spaces, there is v…