3 papers
cs.LG2026
How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness
Polina Gordienko, Georg Schollmeyer, Frauke Kreuter +1
Multi-task benchmarks have become a central pillar of machine learning research, yet their growing influence has incentivised benchmark gaming -- strategic actions taken to improve…
cs.LG2026
Beyond Arrow: From Impossibility to Possibilities in Multi-Criteria Benchmarking
Polina Gordienko, Christoph Jansen, Julian Rodemann +1
Modern benchmarks such as HELM MMLU account for multiple metrics like accuracy, robustness and efficiency. When trying to turn these metrics into a single ranking, natural aggregat…
cs.AI2025
Consensus in Motion: A Case of Dynamic Rationality of Sequential Learning in Probability Aggregation
Polina Gordienko, Christoph Jansen, Thomas Augustin +1
We propose a framework for probability aggregation based on propositional probability logic. Unlike conventional judgment aggregation, which focuses on static rationality, our mode…